EDBT 2026 Demo / reviewers in the wild / expert
Mohammad Azzeh
dblp:61/376
· DBLP profile ↗
35ranked-venue papers
22as first author
17since 2021 · last 2026
0000-0002-0323-6452ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 10 first-author · 8 since 2021Software engineering, systems software and programming languages · 16 · 12 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-project software defects prediction using fuzzy embedding and deep learningabstractContext: Cross-project defect prediction (CPDP) aims to predict software defects in a target project using data from related source projects, especially when defect data for the target project is limited or unavailable. A key challenge in CPDP is data heterogeneity and distributional differences across projects, which often lead to poor performance and unreliable predictions. Objectives: This study proposes a new CPDP model that improves prediction accuracy and robustness by introducing fuzzy embedding and deep learning to better capture similarities and differences across projects. The method is designed to mitigate mismatches in data distribution that hinder existing transfer learning and transformation-based approaches. Methods: The fuzzy embedding technique is built on fuzzy clustering and fuzzy set theory, which map each data point into a two-dimensional space of membership degrees across clusters. This representation models complex relationships with partial memberships and preserves contextual information that is often lost in conventional transformations. A deep learning model based on convolutional neural networks (CNN) processes the embedding matrices to learn discriminative defect patterns. The framework is evaluated against state-of-the-art CPDP models across multiple datasets, and sensitivity analysis is conducted on the number of clusters used in fuzzy embeddings. Results: Empirical evaluation shows that the proposed model consistently outperforms advanced CPDP approaches that rely on standard transformation or weighting methods. Improvements are observed in AUC and other key metrics across diverse datasets, demonstrating that fuzzy embeddings enhance the ability of deep learning to generalize knowledge across projects with varying characteristics. Conclusion: This work contributes a practical and effective solution for addressing heterogeneity in CPDP. By combining fuzzy embeddings with deep learning, the model not only achieves higher predictive accuracy but also improves reliability in real-world scenarios where project data distributions differ significantly. These findings highlight the potential of fuzzy embedding to support more resilient software quality assurance practices and provide actionable insights for practitioners dealing with limited or imbalanced defect data. Mohammad Azzeh, Mohammad Abdel-Rahman |
Inf. Softw. Technol. | 1 |
| 2026 | A Systematic Review on Code Smell Detection Approaches in Open Source ProjectsabstractABSTRACT Background Code smells are an indicator that something is misplaced in software systems, which can reduce maintainability and quality of software applications. In the literature, there is much research about code smells where numerous strategies for automating code smell detection have been developed to improve software quality. Objective The purpose of this work is to provide a review of search‐based, heuristic‐based, machine learning‐based, deep learning‐based, and hybrid‐based code smell detection algorithms. Method Concerning the main goals of this research, we have found 38 primary studies. We gathered relevant studies published on this topic between 2017 and 2024. These articles' data were extracted according to some criteria, including code smells, machine learning methods, programming languages, dataset size, evaluation strategy, and statistical tests. Results Machine learning‐based code smell detection methods have been suggested in most empirical investigations. We found that machine learning and deep learning are the most popular approaches for predicting code smells. Researchers typically employ support vector machine and decision tree algorithms. The Arcelli Fontana and Zanoni benchmark dataset was the most frequently investigated dataset. Most of the research community's attention has been focused on a small number of smells, including blob, feature envy, long method, and data class. Researchers also pay more attention to code smells like Long Method and Feature Envy. Deep learning techniques are increasingly used, and most scientists utilize source code metrics as indicators. The standard performance measurements mostly used are F‐measure, recall, and precision. Conclusions The study provides an overview of existing approaches and highlights current research trends in code smell detection, particularly the increasing use of machine learning and deep learning techniques. Sawsan Alodibat, Mohammad Azzeh |
Softw. Pract. Exp. | 2 |
| 2025 | Correction: Software effort estimation using convolutional neural network and fuzzy clustering
Mohammad Azzeh, Abedalrhman Alkhateeb, Ali Bou Nassif |
Neural Comput. Appl. | 1 |
| 2025 | Integrative multi-omics approach for identifying and classifying high-risk patients in prostate cancer management
Mohammad Azzeh, Noor Afeshat, Hazem Qattous, Abedrahman Alkhateeb |
Neural Comput. Appl. | 1 |
| 2024 | Software effort estimation using convolutional neural network and fuzzy clustering
Mohammad Azzeh, Abedalrhman Alkhateeb, Ali Bou Nassif |
Neural Comput. Appl. | 1 |
| 2024 | Parameter-efficient fine-tuning of pre-trained code models for just-in-time defect prediction
Manar Abu Talib, Ali Bou Nassif, Mohammad Azzeh, Yaser Alesh, Yaman Afadar |
Neural Comput. Appl. | 3 |
| 2024 | A soft computing approach for software defect density predictionabstractAbstract Defect density is an essential software testing and maintenance aspect that determines the quality of software products. It is used as a management factor to distribute limited human resources successfully. The availability of public defect datasets facilitates building defect density prediction models using established static code metrics. Since the data gathered for software modules are often subject to uncertainty, it becomes difficult to deliver accurate and reliable predictions. To alleviate this issue, we propose a new prediction model that integrates gray system theory and fuzzy logic to handle the imprecision in software measurement. We propose a new similarity measure that combines the benefits of fuzzy logic and gray relational analysis. The proposed model was validated against defect density prediction models using public defect datasets. The defect density variable is frequently sparse because of the vast number of none‐defected modules in the datasets. Therefore, we also check our proposed model's performance against the sparsity level. The findings reveal that the developed model surpasses other defect density prediction models over the datasets with high and very high sparsity ratios. The ensemble learning techniques are competitive choices to the proposed model when the sparsity ratio is relatively small. On the other hand, the statistical regression models were the most inadequate methods for such problems and datasets. Finally, the proposed model was evaluated against different degrees of uncertainty using a sensitivity analysis procedure. The results showed that our model behaves stably under different degrees of uncertainty. Mohammad Azzeh, Yousef Alqasrawi, Yousef Elsheikh |
J. Softw. Evol. Process. | 1 |
| 2024 | Software defect density prediction using grey system theory and fuzzy logic
Mohammad Azzeh, Yousef Elsheikh, Yousef Alqasrawi |
Soft Comput. | 1 |
| 2023 | Truth Seeker of the Largest Social Media Content using Machine Learning AlgorithmsabstractSocial media is the source of news, information, and opinions for many people. So, the pursuit of truth in the realm of social media is a noble and essential endeavor because it underpins informed decision-making, counters misinformation, nurtures critical thinking and builds trust. We used machine learning(ML) algorithms and Natural Language Processing(NLP) to understand and process human language used in social media posts of the Largest Social Media Ground-Truth Dataset for Real/Fake Content to create a model that could be used as a reference to detect fake news from Twitter posts, which is a binary classification problem. We created two different datasets from the original large dataset using different NLP vectorization techniques (word2vec and TF -IDF), generating different features per dataset. Then we applied seven ML algorithms on each dataset and compared the results by extracting the confusion matrix and ROC per model. Random forest and AdBoost algorithms achieved the best F1 scores. Maysa Khalil, Mohammad Azzeh |
ICMLA | 2 |
| 2023 | An optimized case-based software project effort estimation using genetic algorithm
Shaima Hameed, Yousef Elsheikh, Mohammad Azzeh |
Inf. Softw. Technol. | 3 |
| 2023 | Examining the performance of kernel methods for software defect prediction based on support vector machine
Mohammad Azzeh, Yousef Elsheikh, Ali Bou Nassif, Lefteris Angelis |
Sci. Comput. Program. | 1 |
| 2022 | Multi-omics Data Integration Model based on Isomap and Convolutional Neural NetworkabstractThe recent advances in genetics technologies have led to high-throughput characterized different biological molecules’ functionalities. The availability of heterogeneous omics sparked the challenge of integrating them for further analysis. This work incorporates the Isomap technique to embed multi-omic data into a convolutional neural network (CNN). The deep learning model fuses three omics data, which are gene expression, copy number alteration (CNA), and DNA methylation data, for breast cancer stage prediction. Isomap is utilized to convert the high-dimensional data into 2-dimensional maps. The gene similarity network (GSN) map is created based on gene expression data to preserve gene relationships. The values from three omics for each sample are used to color the GSN map based on the RGB system. The created GSN maps for all samples are fed into the CNN for classification.The model was applied to TCGA breast invasive carcinoma data set to predict the stage of breast cancer. It outperformed the state-of-art iSOM-GSN model in performance metrics, including accuracy, precision, recall, f1-measure, and area under the curve (AUC). The results indicate that a combination of Isomap embedding technique and CNN can successfully integrate a multi-omics data set for cancer outcome prediction, including the diagnosis and prognosis of the complex disease. Abedalrhman Alkhateeb, Bashier ElKarami, Hazem Qattous, Abdullah Al-Refai, Noor AlAfeshat, Behnam Shahrrava, Mohammad Azzeh |
ICMLA | 7 |
| 2022 | On Selection of Optimal Kernel Function for Software Defect PredictionabstractSoftware defect prediction is used to identify the most likely defect-prone modules. This activity can help the quality assurance team to distribute limited resources efficiently as well as reduce the time and cost of the testing phase. Support Vector Machine (SVM) has been widely used to build software defect prediction models. Nevertheless, the accuracy of such method depends on multiple factors such as the type of SVM method, choice of kernel functions and dimensionality of the dataset. This paper continues along that research line and examines the impact and stability of four kernel functions and feature dimensionality reduction on the performance of SVM for software defect pre- diction. Comprehensive experiments have been conducted using four kernel functions, ten feature selection thresholds based on the Information gain algorithm, and seven public datasets. This has resulted in 280 experiments. The findings demonstrate that there is no kernel function that can show stable performance across different experimental settings. However, both Radial basis and Sigmoid kernel functions showed more stability than Polynomial and Linear kernel functions. In addition, changing the feature set from high to low dimensionality for each kernel function did not always improve the accuracy of SVM, which contradicts the findings that say using a small set of features can be as effective as using a large set of features. Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
ICMLA | 1 |
| 2022 | Locally weighted regression with different kernel smoothers for software effort estimation
Yousef Alqasrawi, Mohammad Azzeh, Yousef Elsheikh |
Sci. Comput. Program. | 2 |
| 2022 | On the value of project productivity for early effort estimation
Mohammad Azzeh, Ali Bou Nassif, Yousef Elsheikh, Lefteris Angelis |
Sci. Comput. Program. | 1 |
| 2021 | Predicting software effort from use case points: A systematic review
Mohammad Azzeh, Ali Bou Nassif, Imtinan Basem Attili |
Sci. Comput. Program. | 1 |
| 2021 | Empirical analysis on productivity prediction and locality for use case points method
Mohammad Azzeh, Ali Bou Nassif, Cuauhtémoc López Martín |
Softw. Qual. J. | 1 |
| 2020 | Transformed k-nearest neighborhood output distance minimization for predicting the defect density of software projects
Cuauhtémoc López Martín, Yenny Villuendas-Rey, Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
J. Syst. Softw. | 3 |
| 2018 | Ensemble of Learning Project Productivity in Software Effort Based on Use Case PointsabstractIt is well recognized that the project productivity is a key driver in estimating software project effort from Use Case Point size metric at early software development stages. Although, there are few proposed models for predicting productivity, there is no consistent conclusion regarding which model is the superior. Therefore, instead of building a new productivity prediction model, this paper presents a new ensemble construction mechanism applied for software project productivity prediction. Ensemble is an effective technique when performance of base models is poor. We proposed a weighted mean method to aggregate predicted productivities based on average of errors produced by training model. The obtained results show that the using ensemble is a good alternative approach when accuracies of base models are not consistently accurate over different datasets, and when models behave diversely. Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan, Cuauhtémoc López Martín |
ICMLA | 1 |
| 2018 | Upsilon-SVR Polynomial Kernel for Predicting the Defect Density in New Software ProjectsabstractAn important product measure to determine the effectiveness of software processes is the defect density (DD). In this study, we propose the application of support vector regression (SVR) to predict the DD of new software projects obtained from the International Software Benchmarking Standards Group (ISBSG) Release 2018 data set. Two types of SVR (i.e., ε-SVR and υ-SVR) were applied to train and test these projects. Each SVR used four types of kernels. The prediction accuracy of each SVR was compared to that of a statistical regression (i.e., a simple linear regression, SLR). Statistical significance test showed that υ-SVR with polynomial kernel was better than that of SLR when new software projects were developed on mainframes and coded in programming languages of third generation Cuauhtémoc López Martín, Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
ICMLA | 2 |
| 2018 | Project productivity evaluation in early software effort estimationabstractAbstract The productivity factor has long been a key driver to estimate effort from Use Case Points (UCP) size measure, especially when historical dataset is absent. But, no one questions: Does productivity still matter when historical data are also available? To facilitate answering this question, the present paper studies the role of productivity from 2 perspectives. First, does learning productivity from historical data lead to better accuracy than using fixed productivity ratios? Second, what is the impact of ignoring productivity when estimating the effort from UCP? Five different models that use productivity factor have been used under different experimental settings and compared with some regression models that use only UCP size metrics. We found that dynamically learning and adjusting productivity from historical data are more efficient than using fixed productivity values. Moreover, using UCP size variables to estimate effort tends to be more accurate than using productivity and UCP variables. We also did not find any significant improvement when using UCP adjustment factors for measuring productivity. Finally, we conclude that the productivity factor is a good driver to generate effort estimate from UCP in the presence and absence of historical datasets. But using UCP size variables alone for predicting effort is more accurate than using productivity. Mohammad Azzeh, Ali Bou Nassif |
J. Softw. Evol. Process. | 1 |
| 2017 | Analyzing the relationship between project productivity and environment factors in the use case points methodabstractProject productivity is a key factor for producing effort estimates from use case points (UCP), especially when the historical dataset is absent. The first versions of UCP effort estimation models used a fixed number or very limited numbers of productivity ratios for all new projects. These approaches have not been well examined over a large number of projects, so the validity of these studies was a matter for criticism. The newly available large software datasets allow us to perform further research on the usefulness of productivity for effort estimation of software development. Specifically, we studied the relationship between project productivity and UCP environmental factors, as they have a significant impact on the amount of productivity needed for a software project. Therefore, we designed 4 studies, using various classification and regression methods, to examine the usefulness of that relationship and its impact on UCP effort estimation. The results we obtained are encouraging and show potential improvement in effort estimation. Furthermore, the efficiency of that relationship is better over a dataset that comes from industry because of the quality of data collection. Our comment on the findings is that it is better to exclude environmental factors from calculating UCP and make them available only for computing productivity. The study also encourages project managers to understand how to better assess the environmental factors, as they do have a significant impact on productivity. Mohammad Azzeh, Ali Bou Nassif |
J. Softw. Evol. Process. | 1 |
| 2016 | User Movement Prediction: The Contribution of Machine Learning TechniquesabstractAmbient Assisted Living (AAL) aims to increase the time older people or disabled people can live in their home environment by assisting them in performing activities of daily living by the use of intelligent products. Localization and tracking of users in indoor environment are the main components of AAL. Wireless sensor networks is an effective technology to accomplish these services by using Received Signal Strength (RSS) information. This work seeks to investigate the effect of machine learning techniques on the accuracy of user movement prediction. Five base classifiers and two ensemble learning approaches are employed and the results are evaluated in terms of precision recall, and F-measure. A real-life benchmark dataset in the area of AAL is used for evaluation. The results show that J48 is the best performing model compared to the other base-level classifiers. It also shows that Bagged J48 achieves the best performance. Shadi Banitaan, Mohammad Azzeh, Ali Bou Nassif |
ICMLA | 2 |
| 2016 | Pareto efficient multi-objective optimization for local tuning of analogy-based estimation
Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan, Fadi Almasalha |
Neural Comput. Appl. | 1 |
| 2016 | Guest editorial: special issue on predictive analytics using machine learning
Ali Bou Nassif, Mohammad Azzeh, Shadi Banitaan, Daniel Neagu |
Neural Comput. Appl. | 2 |
| 2016 | Neural network models for software development effort estimation: a comparative study
Ali Bou Nassif, Mohammad Azzeh, Luiz Fernando Capretz, Danny Ho |
Neural Comput. Appl. | 2 |
| 2015 | An Application of Classification and Class Decomposition to Use Case Point Estimation MethodabstractUse Case Points (UCP) estimation method describes the process of computing the software project size and productivity from use case diagram elements. These metrics are then used to predict the project effort at early stage of software development. The main challenges with previous models are that they were constructed based on a very limited number of observations, and using limited productivity ratios. This paper presents a new approach to predict productivity from UCP environmental factors by applying classification with decomposition technique. A class decomposition provides a number of advantages to supervised learning algorithms through segmenting classes into more homogenous classes, and therefore, increase their diversity. The proposed model is constructed and validated over two datasets that have relatively sufficient number of observations. The accuracy results are promising and have potential to increase accuracy of early effort estimation. Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
ICMLA | 1 |
| 2015 | Class Decomposition Using K-Means and Hierarchical ClusteringabstractThis paper presents a clustering-based class decomposition approach to improve the performance of classifiers. Class decomposition works by dividing each class into clusters, and by relabeling instances contained by each cluster with a new class. Several case studies used class decomposition combined with linear classifiers. While there is an essential improvement in classification accuracy because of class decomposition, the most effective clustering algorithm is not obvious. The aim of this work is to investigate the effect of two clustering algorithms, K-means and hierarchical, on class decomposition. In this work, we study class decomposition when combined with the Naive Bayes classifier using four real-world datasets. Experimental results show an improvement in classification accuracy for most of the datasets when class decomposition using both K-means and hierarchical clustering is performed. The results also show that class decomposition is not suitable for all datasets. Shadi Banitaan, Ali Bou Nassif, Mohammad Azzeh |
ICMLA | 3 |
| 2015 | An empirical evaluation of ensemble adjustment methods for analogy-based effort estimation
Mohammad Azzeh, Ali Bou Nassif, Leandro L. Minku |
J. Syst. Softw. | 1 |
| 2013 | Fuzzy Model Tree for Early Effort EstimationabstractUse Case Points (UCP) is a well-known method to estimate the project size, based on Use Case diagram, at early phases of software development. Although the Use Case diagram is widely accepted as a de-facto model for analyzing object oriented software requirements over the world, UCP method did not take sufficient amount of attention because, as yet, there is no consensus on how to produce software effort from UCP. This paper aims to study the potential of using Fuzzy Model Tree to derive effort estimates based on UCP size measure using a dataset collected for that purpose. The proposed approach has been validated against Tree boost model, Multiple Linear Regression and classical effort estimation based on the UCP model. The obtained results are promising and show better performance than those obtained by classical UCP, Multiple Linear Regression and slightly better than those obtained by Tree boost model. Mohammad Azzeh, Ali Bou Nassif |
ICMLA (2) | 1 |
| 2012 | A Treeboost Model for Software Effort Estimation Based on Use Case PointsabstractSoftware effort prediction is an important task in the software development life cycle. Many models including regression models, machine learning models, algorithmic models, expert judgment and estimation by analogy have been widely used to estimate software effort and cost. In this work, a Tree boost (Stochastic Gradient Boosting) model is put forward to predict software effort based on the Use Case Point method. The inputs of the model include software size in use case points, productivity and complexity. A multiple linear regression model was created and the Tree boost model was evaluated against the multiple linear regression model, as well as the use case point model by using four performance criteria: MMRE, PRED, MdMRE and MSE. Experiments show that the Tree boost model can be used with promising results to estimate software effort. Ali Bou Nassif, Luiz Fernando Capretz, Danny Ho, Mohammad Azzeh |
ICMLA (2) | 4 |
| 2012 | A replicated assessment and comparison of adaptation techniques for analogy-based effort estimation
Mohammad Azzeh |
Empir. Softw. Eng. | 1 |
| 2011 | Adjusted Case-Based Software Effort Estimation Using Bees Optimization Algorithm
Mohammad Azzeh |
KES (2) | 1 |
| 2011 | Analogy-based software effort estimation using Fuzzy numbers
Mohammad Azzeh, Daniel Neagu, Peter I. Cowling |
J. Syst. Softw. | 1 |
| 2010 | Fuzzy grey relational analysis for software effort estimation
Mohammad Azzeh, Daniel Neagu, Peter I. Cowling |
Empir. Softw. Eng. | 1 |