Lukasz Radlinski

dblp:32/6177 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
3since 2021 · last 2026
0000-0003-1007-6597ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking data selection strategies for more accurate software effort prediction using the ISBSG dataset
abstract
A common practice in studies on software development effort prediction involving the ISBSG dataset is upfront data selection by keeping only high-quality cases according to the Data Quality Rating and UFP Rating and the narrow group of predictor attributes having no or very few missing values. Hence, a substantial part of the dataset is discarded. This study investigates the impact of training data quality on the performance of models for software effort prediction. It explores whether less restrictive data selection improves predictive accuracy. Model performance was evaluated with a standardised accuracy, a “win–tie–loss” approach, and a matched-pairs rank biserial correlation coefficient. The non-parametric Scott-Knott effect size difference test provided rankings of data selection strategies and prediction techniques. Using a larger training subset, i.e., including more cases despite a small fraction with low ratings and more predictor attributes despite some of them even with up to 80% of missing values, not only did not degrade model performance but, on the contrary, in many cases, improved it. For most models, the larger size of the training dataset was more beneficial than the smaller one, with only high-quality data. Hence, if predictive accuracy is the priority, training models on all cases (even with low ratings) and a broad set of attributes (up to 70%–80% of missing values) is justified. The best-performing prediction techniques were SVM, XGBoost, and neural networks, depending on the data selection strategy. The provided rankings offer helpful guidance for setting up future effort prediction studies.
Lukasz Radlinski, Jakub Swacha
J. Syst. Softw.1
2024 The Trade-off Between Data Volume and Quality in Predicting User Satisfaction in Software Projects
abstract
Most predictive studies involving the ISBSG dataset used only high-quality cases according to the Data Quality Rating and UFP Rating and a few predictors with no or very few missing values. This study investigated the trade-off between data volume and quality when predicting user satisfaction in software projects. Specifically, it explored whether machine learning models would perform better when trained using a larger dataset containing some portion of low-quality data, a smaller dataset with only high-quality data, or an intermediate setting. A standardised accuracy, a “win-tie-loss” approach, and a matched-pairs rank biserial correlation coefficient were used to evaluate predictive performance. The rankings of data selection strategies for particular models were created using the Scott-Knott Effect Size Difference test. The robustness of results was assessed using Kendall W. For most models, a higher predictive accuracy was achieved when trained on a larger subset, even though it contained some low-quality data. For most models, data selection strategies were robust to data splits. The ranks of data selection strategies were stable across models. Hence, a practical recommendation for predicting user satisfaction, especially when a dataset is small, is to train predictive models on a relatively high-volume subset despite some low-quality data. Provided rankings may be helpful when setting up future experiments on user satisfaction with the ISBSG dataset.
Lukasz Radlinski
SEAA1
2021 Analysis of factors of software development effort and productivity
abstract
The goal of this paper was to identify factors of development effort and productivity and investigate the nature of these relationships using the current release of the ISBSG dataset. In particular, statistical measures of correlations and associations, single-predictor linear regression models, and, most importantly, a moderation analysis were used. Performed analysis demonstrated which attributes are in strong relationships with effort and productivity, investigated the explainability of single-predictor models, and discussed if and how particular attributes moderate the strongest relationships reflected in these single-predictor models.
Lukasz Radlinski
KES1
2020 Predicting User Satisfaction in Software Projects using Machine Learning Techniques
Lukasz Radlinski
ENASE1
2020 Stability of user satisfaction prediction in software projects
abstract
The goal of this paper was to investigate the stability of predictions of user satisfaction using the extended version of the ISBSG dataset. The analysis involved building and training 40 models using 12 machine learning techniques. The results were analysed using a ‘win-tie-loss’ procedure based on a statistical significance. The overall best performing models were random forests. However, 19 models using ten techniques performed the best in at least one of 20 passes. High variability of data across passes caused the difficulty of predictions in some passes.
Lukasz Radlinski
KES1
2016 From complex questionnaire and interviewing data to intelligent Bayesian network models for medical decision support
Anthony C. Constantinou, Norman E. Fenton, William Marsh 0001, Lukasz Radlinski
Artif. Intell. Medicine4
2013 Predicting the Flow of Defect Correction Effort using a Bayesian Network Model
Thomas Schulz, Lukasz Radlinski, Thomas Gorges, Wolfgang Rosenstiel
Empir. Softw. Eng.2
2012 Empirical Analysis of the Impact of Requirements Engineering on Software Quality
Lukasz Radlinski
REFSQ1
2011 A Framework for Integrated Software Quality Prediction Using Bayesian Nets
Lukasz Radlinski
ICCSA (5)1
2010 Software Development Effort and Quality Prediction Using Bayesian Nets and small Local Qualitative Data
Lukasz Radlinski
SEKE1
2008 On the effectiveness of early life cycle defect prediction with Bayesian Nets
Norman E. Fenton, Martin Neil, William Marsh 0001, Peter Stewart Hearty, Lukasz Radlinski, Paul Krause
Empir. Softw. Eng.5