VLDB 2026 Research / reviewers in the wild / expert
Martin J. Shepperd
dblp:08/3451
· DBLP profile ↗
78ranked-venue papers
24as first author
6since 2021 · last 2026
0000-0003-1874-6145ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 67 · 22 first-author · 6 since 2021Artificial intelligence and machine learning · 7Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An audit of machine learning experiments on software defect predictionabstractMachine learning algorithms are increasingly being proposed to solve the problem of predicting defect-prone software components. In this literature, computational experiments are the primary means of evaluating and comparing learners and the credibility of findings depends critically on their experimental design and reporting. This paper audits recent software defect prediction (SDP) experiments by assessing their experimental design, analysis and reporting practices against widely accepted norms from statistics, machine learning and empirical software engineering. Our aim is to characterise the current state of practice and evaluate the reproducibility of published findings. We undertook an audit of relevant studies published from the SCOPUS database (2019-2023) focusing on their experimental design and analysis choices e.g., the outcome variables such as F-measure and the type of out of sample (OOS) validation regime, e.g., cross-validation, plus the statistical analysis and inference mechanisms. In all, we evaluated nine different study issues. This was complemented by an assessment of reproducibility using the instrument proposed by González-Barahona and Robles. Our search located approximately 1,585 experiments in SDP (2019-2023), a substantial body of work. From this, we randomly sampled 101 ( $$ \approx 6.4\%$$ ) papers, 61 journal and 40 conference papers. Almost 50% are behind ‘paywalls’. We found considerable divergence in research practice. The number of datasets used ranged 1-365, the number of learners or learner variants evaluated from 1-34 and the number of performance metrics from 1 to 9. Approximately 45% of papers made use of formal statistical inference. We detected a total of 427 issues distributed across 101 papers (median=4) with only one paper being entirely issue-free. In terms of reproducibility, experiments ranged from near perfect to lacking almost all required information. We also found two examples of tortured phrases and potential “paper mill” activity. Approaches to designing and reporting computational experiments varied greatly, but almost half the studies provided insufficient information such that reproduction would be challenging. Overall, our audit suggests that as a research community, we have considerable scope for improvement. Fortunately, many improvements should be neither difficult nor costly to achieve. Giuseppe Destefanis, Leila Yousefi, Martin J. Shepperd, Allan Tucker, Stephen Swift, Steve Counsell, Mahir Arzoky |
Empir. Softw. Eng. | 3 |
| 2026 | LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviewsabstractContext: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matrix metrics used to evaluate them can mislead under the imbalanced, cost-asymmetric conditions of screening. Objective: We develop and justify LLM4SCREENLIT—practical recommendations for researchers conducting LLM-screening evaluations and for editors and reviewers assessing such studies—differentiated by study type (retrospective benchmarking vs. deployment for a specific SR). Method: Using Delgado-Chaves et al. (2025), an 18-LLM benchmark across three biomedical SRs, as a motivating example, we reviewed 28 additional papers and extracted their reported metrics. We propose a Weighted Matthews Correlation Coefficient (WMCC) that integrates MCC’s chance-correction with asymmetric misclassification costs, and validated it on three software-engineering (SE) reanalyses (Felizardo et al. 2024; Syriani et al. 2024; Huotala et al. 2025), the largest covering 9 LLMs × 24 SE secondary studies (34,528 articles). Results: Across the 29 papers, only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost. In the largest SE reanalysis, MCC and WMCC disagree on the best LLM in 55% of evaluable studies; in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%. Sensitivity analysis (median crossover at w ≈ 2 . 7 , all < 7 ) supports w = 10 as a conservative default. Conclusions: SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available. Editors and reviewers should require these elements as routine. Extension to full-text screening and data extraction is principled but pending empirical validation. Lech Madeyski, Barbara A. Kitchenham, Martin J. Shepperd |
Inf. Softw. Technol. | 3 |
| 2025 | "Estimating Software Project Effort Using Analogies": Reflections After 28 YearsabstractThis invited paper is the result of an invitation to write a retrospective article on a “TSE most influential paper” as part of the journal's 50th anniversary. The objective is to reflect on the progress of software engineering prediction research using the lens of a selected, highly cited research paper and 28 years of hindsight. The paper examines (i) what was achieved, (ii) what has endured and (iii) what could have been done differently with the benefit of retrospection. While many specifics of software project effort prediction have evolved, key methodological issues remain relevant. The original study emphasised empirical validation with benchmarks, out-of-sample testing and data/tool sharing. Four areas for improvement are identified: (i) stronger commitment to Open Science principles, (ii) focus on effect sizes and confidence intervals, (iii) reporting variability alongside typical results and (iv) more rigorous examination of threats to validity. Martin J. Shepperd |
IEEE Trans. Software Eng. | 1 |
| 2024 | Improving classifier-based effort-aware software defect prediction by reducing ranking errorsabstractContext: Software defect prediction utilizes historical data to direct software quality assurance resources to potentially problematic components. Effort-aware (EA) defect prediction prioritizes more bug-like components by taking cost-effectiveness into account. In other words, it is a ranking problem, however, existing ranking strategies based on classification, give limited consideration to ranking errors. Objective: Improve the performance of classifier-based EA ranking methods by focusing on ranking errors. Method: We propose a ranking score calculation strategy called EA-Z which sets a lower bound to avoid near-zero ranking errors. We investigate four primary EA ranking strategies with 16 classification learners, and conduct the experiments for EA-Z and the other four existing strategies. Results: Experimental results from 72 data sets show EA-Z is the best ranking score calculation strategy in terms of Recall@20% and Popt when considering all 16 learners. For particular learners, imbalanced ensemble learner UBag-svm and UBst-rf achieve top performance with EA-Z. Conclusion: Our study indicates the effectiveness of reducing ranking errors for classifier-based effort-aware defect prediction. We recommend using EA-Z with imbalanced ensemble learning. Martin J. Shepperd, Ning Li 0022 |
EASE | 2 |
| 2022 | Special issue on information systems quality management in practice
Martin J. Shepperd, Fernando Brito e Abreu, Ricardo Pérez-Castillo |
Softw. Qual. J. | 1 |
| 2021 | The impact of using biased performance metrics on software defect prediction research
Jingxiu Yao, Martin J. Shepperd |
Inf. Softw. Technol. | 2 |
| 2020 | Using the Lexicon from Source Code to Determine Application DomainabstractContext: The vast majority of software engineering research is reported independently of the application domain: techniques and tools usage is reported without any domain context. As reported in previous research, this has not always been so: early in the computing era, the research focus was frequently application domain specific (for example, scientific and data processing). Andrea Capiluppi, Nemitari Ajienka, Nour Ali, Mahir Arzoky, Steve Counsell, Giuseppe Destefanis, Alina Dana Miron, Bhaveet Nagaria, Rumyana Neykova, Martin J. Shepperd, Stephen Swift, Allan Tucker |
EASE | 10 |
| 2020 | Reasoning about Uncertainty in Empirical ResultsabstractConclusions that are drawn from experiments are subject to varying degrees of uncertainty. For example, they might rely on small data sets, employ statistical techniques that make assumptions that are hard to verify, or there may be unknown confounding factors. In this paper we propose an alternative but complementary mechanism to explicitly incorporate these various sources of uncertainty into reasoning about empirical findings, by applying Subjective Logic. To do this we show how typical traditional results can be encoded as "subjective opinions" -- the building blocks of Subjective Logic. We demonstrate the value of the approach by using Subjective Logic to aggregate empirical results from two large published studies that explore the relationship between programming languages and defects or failures. Neil Walkinshaw, Martin J. Shepperd |
EASE | 2 |
| 2020 | Assessing software defection prediction performance: why using the Matthews correlation coefficient mattersabstractContext: There is considerable diversity in the range and design of computational experiments to assess classifiers for software defect prediction. This is particularly so, regarding the choice of classifier performance metrics. Unfortunately some widely used metrics are known to be biased, in particular F1. Jingxiu Yao, Martin J. Shepperd |
EASE | 2 |
| 2020 | A systematic review of unsupervised learning techniques for software defect prediction
Ning Li 0022, Martin J. Shepperd |
Inf. Softw. Technol. | 2 |
| 2019 | The Prevalence of Errors in Machine Learning Experiments
Martin J. Shepperd, Ning Li 0022, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi |
IDEAL (1) | 1 |
| 2019 | A novel aggregation-based dominance for Pareto-based evolutionary algorithms to configure software product lines
Yani Xue, Miqing Li, Martin J. Shepperd, Stasha Lauria, Xiaohui Liu 0001 |
Neurocomputing | 3 |
| 2019 | "Bad smells" in software analytics papers
Tim Menzies, Martin J. Shepperd |
Inf. Softw. Technol. | 2 |
| 2019 | A Comprehensive Investigation of the Role of Imbalanced Learning for Software Defect PredictionabstractContext: Software defect prediction (SDP) is an important challenge in the field of software engineering, hence much research work has been conducted, most notably through the use of machine learning algorithms. However, class-imbalance typified by few defective components and many non-defective ones is a common occurrence causing difficulties for these methods. Imbalanced learning aims to deal with this problem and has recently been deployed by some researchers, unfortunately with inconsistent results. Objective: We conduct a comprehensive experiment to explore (a) the basic characteristics of this problem; (b) the effect of imbalanced learning and its interactions with (i) data imbalance, (ii) type of classifier, (iii) input metrics and (iv) imbalanced learning method. Method: We systematically evaluate 27 data sets, 7 classifiers, 7 types of input metrics and 17 imbalanced learning methods (including doing nothing) using an experimental design that enables exploration of interactions between these factors and individual imbalanced learning algorithms. This yields 27 × 7 × 7 × 17 = 22491 results. The Matthews correlation coefficient (MCC) is used as an unbiased performance measure (unlike the more widely used F1 and AUC measures). Results: (a) we found a large majority (87 percent) of 106 public domain data sets exhibit moderate or low level of imbalance (imbalance ratio <; 10; median = 3.94); (b) anything other than low levels of imbalance clearly harm the performance of traditional learning for SDP; (c) imbalanced learning is more effective on the data sets with moderate or higher imbalance, however negative results are always possible; (d) type of classifier has most impact on the improvement in classification performance followed by the imbalanced learning method itself. Type of input metrics is not influential. (e) only 52% of the combinations of Imbalanced Learner and Classifier have a significant positive effect. Conclusion: This paper offers two practical guidelines. First, imbalanced learning should only be considered for moderate or highly imbalanced SDP data sets. Second, the appropriate combination of imbalanced method and classifier needs to be carefully chosen to ameliorate the imbalanced learning problem for SDP. In contrast, the indiscriminate application of imbalanced learning can be harmful. Qinbao Song, Martin J. Shepperd |
IEEE Trans. Software Eng. | 3 |
| 2018 | Four commentaries on the use of students and professionals in empirical software engineering experiments
Robert Feldt, Thomas Zimmermann 0001, Gunnar R. Bergersen, Davide Falessi, Andreas Jedlitschka, Natalia Juristo Juzgado, Jürgen Münch, Markku Oivo, Per Runeson, Martin J. Shepperd, Dag I. K. Sjøberg, Burak Turhan |
Empir. Softw. Eng. | 10 |
| 2018 | The role and value of replication in empirical software engineering results
Martin J. Shepperd, Nemitari Ajienka, Steve Counsell |
Inf. Softw. Technol. | 1 |
| 2018 | Authors' Reply to "Comments on 'Researcher Bias: The Use of Machine Learning in Software Defect Prediction'"abstractIn 2014 we published a meta-analysis of software defect prediction studies [1] . This suggested that the most important factor in determining results was Research Group, i.e., who conducts the experiment is more important than the classifier algorithms being investigated. A recent re-analysis [2] sought to argue that the effect is less strong than originally claimed since there is a relationship between Research Group and Dataset. In this response we show (i) the re-analysis is based on a small (21 percent) subset of our original data, (ii) using the same re-analysis approach with a larger subset shows that Research Group is more important than type of Classifier and (iii) however the data are analysed there is compelling evidence that who conducts the research has an effect on the results. This means that the problem of researcher bias remains. Addressing it should be seen as a matter of priority amongst those of us who conduct and publish experiments comparing the performance of competing software defect prediction systems. Martin J. Shepperd, Tracy Hall, David Bowes |
IEEE Trans. Software Eng. | 1 |
| 2016 | Realistic assessment of software effort estimation modelsabstractContext: It is unclear that current approaches to evaluating or comparing competing software cost or effort models give a realistic picture of how they would perform in actual use. Specifically, we're concerned that the usual practice of using all data with some holdout strategy is at variance with the reality of a data set growing as projects complete. Boyce Sigweni, Martin J. Shepperd, Tommaso Turchi |
EASE | 2 |
| 2016 | An External Replication on the Effects of Test-driven Development Using a Multi-site Blind Analysis ApproachabstractContext: Test-driven development (TDD) is an agile practice claimed to improve the quality of a software product, as well as the productivity of its developers. A previous study (i.e., baseline experiment) at the University of Oulu (Finland) compared TDD to a test-last development (TLD) approach through a randomized controlled trial. The results failed to support the claims. Goal: We want to validate the original study results by replicating it at the University of Basilicata (Italy), using a different design. Method: We replicated the baseline experiment, using a crossover design, with 21 graduate students. We kept the settings and context as close as possible to the baseline experiment. In order to limit researchers bias, we involved two other sites (UPM, Spain, and Brunel, UK) to conduct blind analysis of the data. Results: The Kruskal-Wallis tests did not show any significant difference between TDD and TLD in terms of testing effort (p-value = .27), external code quality (p-value = .82), and developers' productivity (p-value = .83). Nevertheless, our data revealed a difference based on the order in which TDD and TLD were applied, though no carry over effect. Conclusions: We verify the baseline study results, yet our results raises concerns regarding the selection of experimental objects, particularly with respect to their interaction with the order in which of treatments are applied. Davide Fucci, Giuseppe Scanniello, Simone Romano 0001, Martin J. Shepperd, Boyce Sigweni, Fernando Uyaguari, Burak Turhan, Natalia Juristo Juzgado, Markku Oivo |
ESEM | 4 |
| 2015 | Using blind analysis for software engineering experimentsabstractContext: In recent years there has been growing concern about conflicting experimental results in empirical software engineering. This has been paralleled by awareness of how bias can impact research results. Boyce Sigweni, Martin J. Shepperd |
EASE | 2 |
| 2014 | Researcher Bias: The Use of Machine Learning in Software Defect PredictionabstractBackground. The ability to predict defect-prone software components would be valuable. Consequently, there have been many empirical studies to evaluate the performance of different techniques endeavouring to accomplish this effectively. However no one technique dominates and so designing a reliable defect prediction model remains problematic. Objective. We seek to make sense of the many conflicting experimental results and understand which factors have the largest effect on predictive performance. Method. We conduct a meta-analysis of all relevant, high quality primary studies of defect prediction to determine what factors influence predictive performance. This is based on 42 primary studies that satisfy our inclusion criteria that collectively report 600 sets of empirical prediction results. By reverse engineering a common response variable we build a random effects ANOVA model to examine the relative contribution of four model building factors (classifier, data set, input metrics and researcher group) to model prediction performance. Results. Surprisingly we find that the choice of classifier has little impact upon performance (1.3 percent) and in contrast the major (31 percent) explanatory factor is the researcher group. It matters more who does the work than what is done. Conclusion. To overcome this high level of researcher bias, defect prediction researchers should (i) conduct blind analysis, (ii) improve reporting protocols and (iii) conduct more intergroup studies in order to alleviate expertise issues. Lastly, research is required to determine whether this bias is prevalent in other applications domains. Martin J. Shepperd, David Bowes, Tracy Hall |
IEEE Trans. Software Eng. | 1 |
| 2013 | Guest Editorial for Special Section from Empirical Software Engineering & Measurement (ESEM) 2011
Martin J. Shepperd, Forrest Shull |
Inf. Softw. Technol. | 1 |
| 2013 | Data Quality: Some Comments on the NASA Software Defect DatasetsabstractBackground--Self-evidently empirical analyses rely upon the quality of their data. Likewise, replications rely upon accurate reporting and using the same rather than similar versions of datasets. In recent years, there has been much interest in using machine learners to classify software modules into defect-prone and not defect-prone categories. The publicly available NASA datasets have been extensively used as part of this research. Objective--This short note investigates the extent to which published analyses based on the NASA defect datasets are meaningful and comparable. Method--We analyze the five studies published in the IEEE Transactions on Software Engineering since 2007 that have utilized these datasets and compare the two versions of the datasets currently in use. Results--We find important differences between the two versions of the datasets, implausible values in one dataset and generally insufficient detail documented on dataset preprocessing. Conclusions--It is recommended that researchers 1) indicate the provenance of the datasets they use, 2) report any preprocessing in sufficient detail to enable meaningful replication, and 3) invest effort in understanding the data prior to applying machine learners. Martin J. Shepperd, Qinbao Song, Zhongbin Sun, Carolyn Mair |
IEEE Trans. Software Eng. | 1 |
| 2012 | Special issue on repeatable results in software engineering prediction
Tim Menzies, Martin J. Shepperd |
Empir. Softw. Eng. | 2 |
| 2012 | Evaluating prediction systems in software project estimation
Martin J. Shepperd, Stephen G. MacDonell |
Inf. Softw. Technol. | 1 |
| 2011 | Group project work from the outset: An in-depth teaching experience reportabstractWe redesigned our undergraduate computing programmes to address problems of motivation and outdated content. The primary vehicle for the new curriculum was the group project which formed a central spine for the entire degree right from the first year. In terms of results, thus far this programme has been successfully run once. Failures, drop outs and students required to retake modules have been halved (from an average of 21.6% from the previous 4 years to 9.5%) and students obtaining the top two grades have increased from 25.2% to 38.9%. Whilst we cannot be certain that all improvement is due to the group projects, informally the change has been well received, however, we are looking for areas to improve including the possibility of more structured support for student metacognitive awareness. Martin J. Shepperd |
CSEE&T | 1 |
| 2011 | Predicting software project effort: A grey relational analysis based method
Qinbao Song, Martin J. Shepperd |
Expert Syst. Appl. | 2 |
| 2011 | A General Software Defect-Proneness Prediction FrameworkabstractBACKGROUND - Predicting defect-prone software components is an economically important activity and so has received a good deal of attention. However, making sense of the many, and sometimes seemingly inconsistent, results is difficult. OBJECTIVE - We propose and evaluate a general framework for software defect prediction that supports 1) unbiased and 2) comprehensive comparison between competing prediction systems. METHOD - The framework is comprised of 1) scheme evaluation and 2) defect prediction components. The scheme evaluation analyzes the prediction performance of competing learning schemes for given historical data sets. The defect predictor builds models according to the evaluated learning scheme and predicts software defects with new data according to the constructed model. In order to demonstrate the performance of the proposed framework, we use both simulation and publicly available software defect data sets. RESULTS - The results show that we should choose different learning schemes for different data sets (i.e., no scheme dominates), that small details in conducting how evaluations are conducted can completely reverse findings, and last, that our proposed framework is more effective and less prone to bias than previous approaches. CONCLUSIONS - Failure to properly or fully evaluate a learning scheme can be misleading; however, these problems may be overcome by our proposed framework. Qinbao Song, Zihan Jia, Martin J. Shepperd, Jin Liu 0016 |
IEEE Trans. Software Eng. | 3 |
| 2010 | Data accumulation and software effort predictionabstractBACKGROUND: In reality project managers are constrained by the incremental nature of data collection. Specifically, project observations are accumulated one project at a time. Likewise within-project data are accumulated one stage or phase at a time. However, empirical researchers have given limited attention to this perspective. Stephen G. MacDonell, Martin J. Shepperd |
ESEM | 2 |
| 2010 | Class movement and re-location: An empirical study of Java inheritance evolution
Emal Nasseri, Steve Counsell, Martin J. Shepperd |
J. Syst. Softw. | 3 |
| 2010 | How Reliable Are Systematic Reviews in Empirical Software Engineering?abstractBACKGROUND-The systematic review is becoming a more commonly employed research instrument in empirical software engineering. Before undue reliance is placed on the outcomes of such reviews it would seem useful to consider the robustness of the approach in this particular research context. OBJECTIVE-The aim of this study is to assess the reliability of systematic reviews as a research instrument. In particular, we wish to investigate the consistency of process and the stability of outcomes. METHOD-We compare the results of two independent reviews undertaken with a common research question. RESULTS-The two reviews find similar answers to the research question, although the means of arriving at those answers vary. CONCLUSIONS-In addressing a well-bounded research question, groups of researchers with similar domain experience can arrive at the same review outcomes, even though they may do so in different ways. This provides evidence that, in this context at least, the systematic review is a robust research method. Stephen G. MacDonell, Martin J. Shepperd, Barbara A. Kitchenham, Emilia Mendes |
IEEE Trans. Software Eng. | 2 |
| 2009 | A Literature Review of Expert Problem Solving using Analogy
Carolyn Mair, Miriam Martincova, Martin J. Shepperd |
EASE | 3 |
| 2009 | The Problem of Labels in E-Assessment of DiagramsabstractIn this article we explore a problematic aspect of automated assessment of diagrams. Diagrams have partial and sometimes inconsistent semantics. Typically much of the meaning of a diagram resides in the labels; however, the choice of labeling is largely unrestricted. This means a correct solution may utilize differing yet semantically equivalent labels to the specimen solution. With human marking this problem can be easily overcome. Unfortunately with e-assessment this is challenging. We empirically explore the scale of the problem of synonyms by analyzing 160 student solutions to a UML task. From this we find that cumulative growth of synonyms only shows a limited tendency to reduce at the margin despite using a range of text processing algorithms such as stemming and auto-correction of spelling errors. This finding has significant implications for the ease in which we may develop future e-assessment systems of diagrams, in that the need for better algorithms for assessing label semantic similarity becomes inescapable. Ambikesh Jayal, Martin J. Shepperd |
ACM J. Educ. Resour. Comput. | 2 |
| 2009 | Integrate the GM(1, 1) and Verhulst Models to Predict Software Stage EffortabstractSoftware effort prediction clearly plays a crucial role in software project management. In keeping with more dynamic approaches to software development, it is not sufficient to only predict the whole-project effort at an early stage. Rather, the project manager must also dynamically predict the effort of different stages or activities during the software development process. This can assist the project manager to reestimate effort and adjust the project plan, thus avoiding effort or schedule overruns. This paper presents a method for software physical time stage-effort prediction based on grey models GM(1,1) and Verhulst. This method establishes models dynamically according to particular types of stage-effort sequences, and can adapt to particular development methodologies automatically by using a novel grey feedback mechanism. We evaluate the proposed method with a large-scale real-world software engineering dataset, and compare it with the linear regression method and the Kalman filter method, revealing that accuracy has been improved by at least 28% and 50%, respectively. The results indicate that the method can be effective and has considerable potential. We believe that stage predictions could be a useful complement to whole-project effort prediction methods. Qinbao Song, Stephen G. MacDonell, Martin J. Shepperd |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 2008 | Can k-NN imputation improve the performance of C4.5 with small software project data sets? A comparative evaluation
Qinbao Song, Martin J. Shepperd |
J. Syst. Softw. | 2 |
| 2007 | Filtering, Robust Filtering, Polishing: Techniques for Addressing Quality in Software DataabstractData quality is an important aspect of empirical analysis. This paper compares three noise handling methods to assess the benefit of identifying and either filtering or editing problematic instances. We compare a 'do nothing' strategy with (i) filtering, (ii) robust filtering and (Hi) filtering followed by polishing. A problem is that it is not possible to determine whether an instance contains noise unless it has implausible values. Since we cannot determine the true overall noise level we use implausible val.ues as a proxy measure. In addition to the ability to identify implausible values, we use another proxy measure, the ability to fit a classification tree to the data. The interpretation is low misclassification rates imply low noise levels. We found that all three of our data quality techniques improve upon the 'do nothing' strategy, also that the filtering and polishing was the most effective technique for dealing with noise since we eliminated the fewest data and had the lowest misclassification rates. Unfortunately the polishing process introduces new implausible values. We believe consideration of data quality is an important aspect of empirical software engineering. We have shown that for one large and complex real world data set automated techniques can help isolate noisy instances and potentially polish the values to produce better quality data for the analyst. However this work is at a preliminary stage and it assumes that the proxy measures of lity are appropriate. Gernot Armin Liebchen, Bhekisipho Twala, Martin J. Shepperd, Michelle Cartwright, Mark Stephens |
ESEM | 3 |
| 2007 | Comparing Local and Global Software Effort Estimation Models - Reflections on a Systematic ReviewabstractThe availability of multi-organisation data sets has made it possible for individual organisations to build and apply management models, even if they do not have data of their own. In the absence of any data this may be a sensible option, driven by necessity. However, if both cross-company (or global) and within-company (or local) data are available, which should be used in preference? Several research papers have addressed this question but without any apparent convergence of results. We conduct a systematic review of empirical studies comparing global and local effort prediction systems. We located 10 relevant studies: 3 supported global models, 2 were equivocal and 5 supported local models. The studies do not have converging results. A contributing factor is that they have utilised different local and global data sets and different experimental designs thus there is substantial heterogeneity. We identify the need for common response variables and for common experimental and reporting protocols. Stephen G. MacDonell, Martin J. Shepperd |
ESEM | 2 |
| 2007 | Most cited journal articles in software engineering
Claes Wohlin, Sebastian G. Elbaum, Martin J. Shepperd |
Inf. Softw. Technol. | 3 |
| 2007 | A new imputation method for small software project data sets
Qinbao Song, Martin J. Shepperd |
J. Syst. Softw. | 2 |
| 2007 | A Systematic Review of Software Development Cost Estimation StudiesabstractThis paper aims to provide a basis for the improvement of software-estimation research through a systematic review of previous work. The review identifies 304 software cost estimation papers in 76 journals and classifies the papers according to research topic, estimation approach, research approach, study context and data set. A Web-based library of these cost estimation papers is provided to ease the identification of relevant estimation research results. The review results combined with other knowledge provide support for recommendations for future software cost estimation research, including: 1) increase the breadth of the search for relevant studies, 2) search manually for relevant papers within a carefully selected set of journals when completeness is essential, 3) conduct more studies on estimation methods commonly used by the software industry, and 4) increase the awareness of how properties of the data sets impact the results when evaluating estimation methods Magne Jørgensen, Martin J. Shepperd |
IEEE Trans. Software Eng. | 2 |
| 2006 | Assessing the Quality and Cleaning of a Software Project Data Set: An Experience ReportabstractOBJECTIVE – The aim is to report upon an assessment of the impact noise has on the predictive accuracy by comparing noise handling techniques. METHOD – We describe the process of cleaning a large software management dataset comprising initially of more than 10,000 projects. The data quality is mainly assessed through feedback from the data provider and manual inspection of the data. Three methods of noise correction (polishing, noise elimination and robust algorithms) are compared with each other assessing their accuracy. The noise detection was undertaken by using a regression tree model. RESULTS – Three noise correction methods are compared and different results in their accuracy where noted. CONCLUSIONS – The results demonstrated that polishing improves classification accuracy compared to noise elimination and robust algorithms approaches. Gernot Armin Liebchen, Martin J. Shepperd, Bhekisipho Twala, Michelle Cartwright |
EASE | 2 |
| 2006 | Ensemble of missing data techniques to improve software prediction accuracyabstractSoftware engineers are commonly faced with the problem of incomplete data. Incomplete data can reduce system performance in terms of predictive accuracy. Unfortunately, rare research has been conducted to systematically explore the impact of missing values, especially from the missing data handling point of view. This has made various missing data techniques (MDTs) less significant. This paper describes a systematic comparison of seven MDTs using eight industrial datasets. Our findings from an empirical evaluation suggest listwise deletion as the least effective technique for handling incomplete data while multiple imputation achieves the highest accuracy rates. We further propose and show how a combination of MDTs by randomizing a decision tree building algorithm leads to a significant improvement in prediction performance for missing values up to 50%. Bhekisipho Twala, Michelle Cartwright, Martin J. Shepperd |
ICSE | 3 |
| 2006 | Software Defect Association Mining and Defect Correction Effort PredictionabstractMuch current software defect prediction work focuses on the number of defects remaining in a software system. In this paper, we present association rule mining based methods to predict defect associations and defect correction effort. This is to help developers detect software defects and assist project managers in allocating testing resources more effectively. We applied the proposed methods to the SEL defect data consisting of more than 200 projects over more than 15 years. The results show that, for defect association prediction, the accuracy is very high and the false-negative rate is very low. Likewise, for the defect correction effort prediction, the accuracy for both defect isolation effort prediction and defect correction effort prediction are also high. We compared the defect correction effort prediction method with other types of methods - PART, C4.5, and Naive Bayes - and show that accuracy has been improved by at least 23 percent. We also evaluated the impact of support and confidence levels on prediction accuracy, false-negative rate, false-positive rate, and the number of rules. We found that higher support and confidence levels may not result in higher prediction accuracy, and a sufficient number of rules is a precondition for high prediction accuracy. Qinbao Song, Martin J. Shepperd, Michelle Cartwright, Carolyn Mair |
IEEE Trans. Software Eng. | 2 |
| 2005 | A Short Note on Safest Default Missingness Mechanism Assumptions
Qinbao Song, Martin J. Shepperd, Michelle Cartwright |
Empir. Softw. Eng. | 2 |
| 2005 | Reliability and Validity in Comparative Studies of Software Prediction ModelsabstractEmpirical studies on software prediction models do not converge with respect to the question "which prediction model is best?" The reason for this lack of convergence is poorly understood. In this simulation study, we have examined a frequently used research procedure comprising three main ingredients: a single data sample, an accuracy indicator, and cross validation. Typically, these empirical studies compare a machine learning model with a regression model. In our study, we use simulation and compare a machine learning and a regression model. The results suggest that it is the research procedure itself that is unreliable. This lack of reliability may strongly contribute to the lack of convergence. Our findings thus cast some doubt on the conclusions of any study of competing software prediction models that used this research procedure as a basis of model comparison. Thus, we need to develop more reliable research procedures before we can have confidence in the conclusions of comparative studies of software prediction models. Ingunn Myrtveit, Erik Stensrud, Martin J. Shepperd |
IEEE Trans. Software Eng. | 3 |
| 2004 | The Evolution of Concurrent Control Software Using Genetic Programming
John K. Hart, Martin J. Shepperd |
EuroGP | 2 |
| 2004 | A controlled experiment investigation of an object-oriented design heuristic for maintainability
Ignatios S. Deligiannis, Ioannis Stamelos, Lefteris Angelis, Manos Roumeliotis, Martin J. Shepperd |
J. Syst. Softw. | 5 |
| 2003 | Using Genetic Programming to Improve Software Effort Estimation Based on General Data Sets
Martin Lefley, Martin J. Shepperd |
GECCO | 2 |
| 2003 | An Empirical Analysis of Linear Adaptation Techniques for Case-Based Prediction
Colin Kirsopp, Emilia Mendes, Rahul Premraj, Martin J. Shepperd |
ICCBR | 4 |
| 2003 | An empirical investigation of an object-oriented design heuristic for maintainability
Ignatios S. Deligiannis, Martin J. Shepperd, Manos Roumeliotis, Ioannis Stamelos |
J. Syst. Softw. | 2 |
| 2003 | Combining techniques to optimize effort predictions in software project management
Stephen G. MacDonell, Martin J. Shepperd |
J. Syst. Softw. | 2 |
| 2002 | Search Heuristics, Case-based Reasoning And Software Project Effort Prediction
Colin Kirsopp, Martin J. Shepperd, John K. Hart |
GECCO | 2 |
| 2002 | A Review of Experimental Investigations into Object-Oriented Technology
Ignatios S. Deligiannis, Martin J. Shepperd, Steve Webster, Manos Roumeliotis |
Empir. Softw. Eng. | 2 |
| 2002 | Editorial
Michael Dyer, Martin J. Shepperd, Claes Wohlin |
Inf. Softw. Technol. | 2 |
| 2001 | Issues on the Effective Use of CBR Technology for Software Project Prediction
Gada F. Kadoda, Michelle Cartwright, Martin J. Shepperd |
ICCBR | 3 |
| 2001 | Editorial Note
Martin J. Shepperd, Michael Dyer |
Inf. Softw. Technol. | 1 |
| 2001 | Predicting with Sparse DataabstractIt is well-known that effective prediction of project cost related factors is an important aspect of software engineering. Unfortunately, despite extensive research over more than 30 years, this remains a significant problem for many practitioners. A major obstacle is the absence of reliable and systematic historic data, yet this is a sine qua non for almost all proposed methods: statistical, machine learning or calibration of existing models. The authors describe our sparse data method (SDM) based upon a pairwise comparison technique and T.L. Saaty's (1980) Analytic Hierarchy Process (AHP). Our minimum data requirement is a single known point. The technique is supported by a software tool known as DataSalvage. We show, for data from two companies, how our approach, based upon expert judgement, adds value to expert judgement by producing significantly more accurate and less biased results. A sensitivity analysis shows that our approach is robust to pairwise comparison errors. We then describe the results of a small usability trial with a practicing project manager. From this empirical work, we conclude that the technique is promising and may help overcome some of the present barriers to effective project prediction. Martin J. Shepperd, Michelle Cartwright |
IEEE Trans. Software Eng. | 1 |
| 2001 | Comparing Software Prediction Techniques Using SimulationabstractThe need for accurate software prediction systems increases as software becomes much larger and more complex. We believe that the underlying characteristics: size, number of features, type of distribution, etc., of the data set influence the choice of the prediction system to be used. For this reason, we would like to control the characteristics of such data sets in order to systematically explore the relationship between accuracy, choice of prediction system, and data set characteristic. It would also be useful to have a large validation data set. Our solution is to simulate data allowing both control and the possibility of large (1000) validation cases. The authors compare four prediction techniques: regression, rule induction, nearest neighbor (a form of case-based reasoning), and neural nets. The results suggest that there are significant differences depending upon the characteristics of the data set. Consequently, researchers should consider prediction context when evaluating competing prediction systems. We observed that the more "messy" the data and the more complex the relationship with the dependent variable, the more variability in the results. In the more complex cases, we observed significantly different results depending upon the particular training set that has been sampled from the underlying data set. However, our most important result is that it is more fruitful to ask which is the best prediction system in a particular context rather than which is the "best" prediction system. Martin J. Shepperd, Gada F. Kadoda |
IEEE Trans. Software Eng. | 1 |
| 2000 | On Building Prediction Systems for Software Engineers
Martin J. Shepperd, Michelle Cartwright, Gada F. Kadoda |
Empir. Softw. Eng. | 1 |
| 2000 | An investigation of machine learning based prediction systems
Carolyn Mair, Gada F. Kadoda, Martin Lefley, Keith Phalp, Chris Schofield, Martin J. Shepperd, Steve Webster |
J. Syst. Softw. | 6 |
| 2000 | Quantitative analysis of static models of processes
Keith Phalp, Martin J. Shepperd |
J. Syst. Softw. | 2 |
| 2000 | An Empirical Investigation of an Object-Oriented Software SystemabstractThe paper describes an empirical investigation into an industrial object oriented (OO) system comprised of 133000 lines of C++. The system was a subsystem of a telecommunications product and was developed using the Shlaer-Mellor method (S. Shlaer and S.J. Mellor, 1988; 1992). From this study, we found that there was little use of OO constructs such as inheritance, and therefore polymorphism. It was also found that there was a significant difference in the defect densities between those classes that participated in inheritance structures and those that did not, with the former being approximately three times more defect-prone. We were able to construct useful prediction systems for size and number of defects based upon simple counts such as the number of states and events per class. Although these prediction systems are only likely to have local significance, there is a more general principle that software developers can consider building their own local prediction systems. Moreover, we believe this is possible, even in the absence of the suites of metrics that have been advocated by researchers into OO technology. As a consequence, measurement technology may be accessible to a wider group of potential users. Michelle Cartwright, Martin J. Shepperd |
IEEE Trans. Software Eng. | 2 |
| 1999 | Perspectives on Information Technology in the New Millennium
Michael Dyer, Martin J. Shepperd |
Inf. Softw. Technol. | 2 |
| 1997 | Workshop Summary: Process Modelling and Empirical Studies of Software EvolutionabstractNo abstract available. Rachel Harrison, Martin J. Shepperd, John W. Daly |
ICSE | 2 |
| 1997 | Process Modelling and Empirical Studies of Software Evolution (PMESSE'97) Workshop Report
Rachel Harrison, Lionel C. Briand, John W. Daly, Marc I. Kellner, David Raffo, Martin J. Shepperd |
Empir. Softw. Eng. | 6 |
| 1997 | Estimating Software Project Effort Using AnalogiesabstractAccurate project effort prediction is an important goal for the software engineering community. To date most work has focused upon building algorithmic models of effort, for example COCOMO. These can be calibrated to local environments. We describe an alternative approach to estimation based upon the use of analogies. The underlying principle is to characterize projects in terms of features (for example, the number of interfaces, the development method or the size of the functional requirements document). Completed projects are stored and then the problem becomes one of finding the most similar projects to the one for which a prediction is required. Similarity is defined as Euclidean distance in n-dimensional space where n is the number of project features. Each dimension is standardized so all dimensions have equal weight. The known effort values of the nearest neighbors to the new project are then used as the basis for the prediction. The process is automated using a PC-based tool known as ANGEL. The method is validated on nine different industrial datasets (a total of 275 projects) and in all cases analogy outperforms algorithmic models based upon stepwise regression. From this work we argue that estimation by analogy is a viable technique that, at the very least, can be used by project managers to complement current estimation techniques. Martin J. Shepperd, Chris Schofield |
IEEE Trans. Software Eng. | 1 |
| 1996 | Effort Estimation Using Analogy
Martin J. Shepperd, Chris Schofield, Barbara A. Kitchenham |
ICSE | 1 |
| 1995 | Editorial
Michael Dyer, Martin J. Shepperd |
Inf. Softw. Technol. | 2 |
| 1995 | Comments on "A Metrics Suite for Object Oriented Design"abstractA suite of object oriented software metrics has recently been proposed by S.R. Chidamber and C.F. Kemerer (see ibid., vol. 20, p. 476-94, 1994). While the authors have taken care to ensure their metrics have a sound measurement theoretical basis, we argue that is premature to begin applying such metrics while there remains uncertainty about the precise definitions of many of the quantities to be observed and their impact upon subsequent indirect metrics. In particular, we show some of the ambiguities associated with the seemingly simple concept of the number of methods per class. The usefulness of the proposed metrics, and others, would be greatly enhanced if clearer guidance concerning their application to specific languages were to be provided. Such empirical considerations are as important as the theoretical issues raised by the authors.> Neville Churcher, Martin J. Shepperd |
IEEE Trans. Software Eng. | 2 |
| 1994 | Editorial
Michael Dyer, Martin J. Shepperd |
Inf. Softw. Technol. | 2 |
| 1994 | A critique of three metrics
Martin J. Shepperd, Darrel C. Ince |
J. Syst. Softw. | 1 |
| 1993 | Practical software metrics for project management and process improvement: R Grady Prentice-Hall (1992) £30.95 282 pp ISBN 0 13 720384 5
Martin J. Shepperd |
Inf. Softw. Technol. | 1 |
| 1992 | First International Conference on the Software Process Redondo Beach, CA, USA 21-22 October 1991
Martin J. Shepperd |
Inf. Softw. Technol. | 1 |
| 1992 | Software engineer's reference book J McDermid (ed) Butterworth-Heinemann (1991) £125 hardback ISBN 0-750-61040-9
Martin J. Shepperd |
Inf. Softw. Technol. | 1 |
| 1992 | Products, processes and metrics
Martin J. Shepperd |
Inf. Softw. Technol. | 1 |
| 1992 | Measurement of structure and size of software designs
Martin J. Shepperd |
Inf. Softw. Technol. | 1 |
| 1991 | Software Metrics in Software Engineering and Artificial IntelligenceabstractThis paper examines the utility of much of the software metrics research that has been carried out in software engineering to problems in artificial intelligence. The paper first reviews the work that has been carried out and then makes a number of suggestions about how it could be transferred to the artificial intelligence arena — applying it in particular to expert system development. Martin J. Shepperd, Darrel C. Ince |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 1991 | Design metrics and software maintainability: An experimental investigationabstractAbstract An empirical study was conducted into the relationship between various design metrics and software maintainability. This was based upon maintenance changes made to four different versions of a project management tool carried out by a total of 60 programmers. The overall conclusion from the investigation, was that accurate prediction of quality characteristics for single maintenance changes is extremely difficult. This is due to the many sources of variation—principally change type and programmer ability. Nevertheless, we show that measures of information flow local to specific modifications are significantly related to error rates, with a 600% greater probability of a residual error as a consequence of a change in a module with a high level of information flow‐based coupling, than a module with a low level of coupling. Furthermore, we show that different types of change reveal marked variations in their relationships with the design metrics. Consequently, we argue that using robust statistical techniques and theoretically well‐founded design metrics, engineering approximations are possible. Martin J. Shepperd, Darrel C. Ince |
J. Softw. Maintenance Res. Pract. | 1 |