Khaled Walid Al-Sabbagh

dblp:251/4681 · DBLP profile ↗
← Back
7ranked-venue papers
7as first author
3since 2021 · last 2026
0000-0003-2571-5099ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 The Impact of Class Noise-handling on the Effectiveness of Machine Learning-based Methods for Build Outcome and Code Change Request Predictions
abstract
Machine learning-based methods are increasingly used to optimize build processes and accelerate the integration of software code. These methods leverage large volumes of historical code changes to train models on predicting and preventing issues in the codebase that could delay code integrations and features delivery to end-users. The objective of this study is to examine the impact of handling class noise present in software code changes collected from Continuous Integration (CI) systems on the predictive performance of machine learning models for predicting the execution outcome of CI builds and negative code reviews. In this study, we conduct a series of computational experiments using data from 110 Java open-source projects, examining the effectiveness of two removal-based statistical techniques - Majority Filter (MF) and Consensus Filter (CF) - and two corrective techniques - Domain Knowledge-based (DB) and CleanLab. Our results show that removal-based techniques significantly improve model predictive performance in both build outcome and negative code review prediction tasks. For build outcome prediction, applying MF increased the F1-score from 82% to 97%, and MCC from 0.13 to 0.58. In negative code review predictions, MF improved the F1-score from 17% to 53%, and MCC from −0.03 to 0.57. The DB technique was effective primarily in the context of code review comments but less so for build outcome predictions. While CleanLab yielded more consistent predictions, its overall impact on model performance was more moderate compared to removal-based techniques. Additionally, our findings show that hyperparameter tuning, applied independently or in combination with CleanLab, can further improve model performance; however, these gains did not surpass those achieved by removal-based techniques alone. We conclude that applying removal-based techniques to the training data of code changes is necessary to improve the prediction of build outcomes and negative code review comments.
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig
ACM Trans. Softw. Eng. Methodol.1
2022 Improving Software Regression Testing Using a Machine Learning-Based Method for Test Type Selection
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig
PROFES1
2022 Improving test case selection by handling class and attribute noise
abstract
Big data and machine learning models have been increasingly used to support software engineering processes and practices. One example is the use of machine learning models to improve test case selection in continuous integration. However, one of the challenges in building such models is the large volume of noise that comes in data, which impedes their predictive performance. In this paper, we address this issue by studying the effect of two types of noise, called class and attribute, on the predictive performance of a test selection model. For this purpose, we analyze the effect of class noise by using an approach that relies on domain knowledge for relabeling contradictory entries and removing duplicate ones. Thereafter, an existing approach from the literature is used to experimentally study the effect of attribute noise removal on learning. The analysis results show that the best learning is achieved when training a model on class-noise cleaned data only — irrespective of attribute noise. Specifically, the learning performance of the model reported 81% precision, 87% recall, and 84% f-score compared with 44% precision, 17% recall, and 25% f-score for a model built on uncleaned data. Finally, no causality relationship between attribute noise removal and the learning of a model for test case selection was drawn.
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig
J. Syst. Softw.1
2020 Improving Data Quality for Regression Test Selection by Reducing Annotation Noise
abstract
Big data and machine learning models have been increasingly used to support software engineering processes and practices. One example is the use of machine learning models to improve test case selection in continuous integration. However, one of the challenges in building such models is the identification and reduction of noise that often comes in large data. In this paper, we present a noise reduction approach that deals with the problem of contradictory training entries. We empirically evaluate the effectiveness of the approach in the context of selective regression testing. For this purpose, we use a curated training set as input to a tree-based machine learning ensemble and compare the classification precision, recall, and f-score against a non-curated set. Our study shows that using the noise reduction approach on the training instances gives better results in prediction with an improvement of 37% on precision, 70% on recall, and 59% on f-score.
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig, Wilhelm Meding
SEAA1
2020 The Effect of Class Noise on Continuous Test Case Selection: A Controlled Experiment on Industrial Data
Khaled Walid Al-Sabbagh, Regina Hebig, Miroslaw Staron
PROFES1
2019 Predicting Test Case Verdicts Using Textual Analysis of Committed Code Churns
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig, Wilhelm Meding
IWSM-Mensura1
2018 The connections between group maturity, software development velocity, and planning effectiveness
abstract
Abstract Empirical evidence regarding the connection between group development (maturity) and the success of software development teams is lacking. The purpose of this research is to gain a qualitative and quantitative understanding of how velocity and planning effectiveness of software teams connect to a group development model. The Group Development Questionnaire was given to 19 software developers from 4 work groups to assess their group development maturity. The work groups' responses to the survey were checked for correlation with development velocity and planning effectiveness. Semistructured interviews were conducted with 16 individuals from the same 4 work groups to explore issues about their group maturity and to validate the responses of the interviewees in the Group Development Questionnaire. The measurement of the fourth stage of group development had a strong association with the planning effectiveness measurement, which means that a team with less issues in the fourth phase of group development is more effective in adhering to its plans. On the other hand, group development and velocity showed no significant convergent validity. We conclude that the dynamics within software development teams might correlate to their ability to deliver the expected outcome as planned but not to their ability to develop tasks faster.
Khaled Walid Al-Sabbagh, Lucas Gren
J. Softw. Evol. Process.1