VLDB 2026 Research / reviewers in the wild / expert
Barbara A. Kitchenham
dblp:k/BarbaraAKitchenham · also Barbara Ann Kitchenham
· DBLP profile ↗
108ranked-venue papers
55as first author
9since 2021 · last 2026
0000-0002-6134-8460ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 107 · 55 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviewsabstractContext: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matrix metrics used to evaluate them can mislead under the imbalanced, cost-asymmetric conditions of screening. Objective: We develop and justify LLM4SCREENLIT—practical recommendations for researchers conducting LLM-screening evaluations and for editors and reviewers assessing such studies—differentiated by study type (retrospective benchmarking vs. deployment for a specific SR). Method: Using Delgado-Chaves et al. (2025), an 18-LLM benchmark across three biomedical SRs, as a motivating example, we reviewed 28 additional papers and extracted their reported metrics. We propose a Weighted Matthews Correlation Coefficient (WMCC) that integrates MCC’s chance-correction with asymmetric misclassification costs, and validated it on three software-engineering (SE) reanalyses (Felizardo et al. 2024; Syriani et al. 2024; Huotala et al. 2025), the largest covering 9 LLMs × 24 SE secondary studies (34,528 articles). Results: Across the 29 papers, only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost. In the largest SE reanalysis, MCC and WMCC disagree on the best LLM in 55% of evaluable studies; in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%. Sensitivity analysis (median crossover at w ≈ 2 . 7 , all < 7 ) supports w = 10 as a conservative default. Conclusions: SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available. Editors and reviewers should require these elements as routine. Extension to full-text screening and data extraction is principled but pending empirical validation. Lech Madeyski, Barbara A. Kitchenham, Martin J. Shepperd |
Inf. Softw. Technol. | 2 |
| 2025 | Using rapid reviews to support software engineering practice: a systematic review and a replication study
Sebastián Pizard, Joaquín Lezama, Rodrigo García, Diego Vallespir, Barbara A. Kitchenham |
Empir. Softw. Eng. | 5 |
| 2025 | Evidence-Based Software Engineering Guidelines RevisitedabstractIn 2002, the authors and their colleagues proposed some preliminary guidelines for empirical software engineering research. In this paper, we revisit them. We believe that for the purpose of supporting the development of project-based bespoke software, they still perform reasonably well. However, the worlds of software and software engineering have changed dramatically in the last 25 years. We suggest that new guidelines are needed to respond to changes in software practices, including the possibility of regulation of AI-based software products, and propose the scope of such guidelines. Shari Lawrence Pfleeger, Barbara A. Kitchenham |
IEEE Trans. Software Eng. | 2 |
| 2024 | Recommendations for analysing and meta-analysing small sample size software engineering experimentsabstractAbstract Context Software engineering (SE) experiments often have small sample sizes. This can result in data sets with non-normal characteristics, which poses problems as standard parametric meta-analysis, using the standardized mean difference (StdMD) effect size, assumes normally distributed sample data. Small sample sizes and non-normal data set characteristics can also lead to unreliable estimates of parametric effect sizes. Meta-analysis is even more complicated if experiments use complex experimental designs, such as two-group and four-group cross-over designs, which are popular in SE experiments. Objective Our objective was to develop a validated and robust meta-analysis method that can help to address the problems of small sample sizes and complex experimental designs without relying upon data samples being normally distributed. Method To illustrate the challenges, we used real SE data sets. We built upon previous research and developed a robust meta-analysis method able to deal with challenges typical for SE experiments. We validated our method via simulations comparing StdMD with two robust alternatives: the probability of superiority ( $$\hat{p}$$ p ^ ) and Cliffs’ d. Results We confirmed that many SE data sets are small and that small experiments run the risk of exhibiting non-normal properties, which can cause problems for analysing families of experiments. For simulations of individual experiments and meta-analyses of families of experiments, $$\hat{p}$$ p ^ and Cliff’s d consistently outperformed StdMD in terms of negligible small sample bias. They also had better power for log-normal and Laplace samples, although lower power for normal and gamma samples. Tests based on $$\hat{p}$$ p ^ always had better or equal power than tests based on Cliff’s d, and across all but one simulation condition, $$\hat{p}$$ p ^ Type 1 error rates were less biased. Conclusions Using $$\hat{p}$$ p ^ is a low-risk option for analysing and meta-analysing data from small sample-size SE randomized experiments. Parametric methods are only preferable if you have prior knowledge of the data distribution. Barbara A. Kitchenham, Lech Madeyski |
Empir. Softw. Eng. | 1 |
| 2023 | Assessing attitudes towards evidence-based software engineering in a government agency
Sebastián Pizard, Fernando Acerenza, Diego Vallespir, Barbara A. Kitchenham |
Inf. Softw. Technol. | 4 |
| 2023 | How Should Software Engineering Secondary Studies Include Grey Material?abstractContext: Recent papers have proposed the use ofgrey literature(GL) and multivocal reviews. These papers have raised issues about the practices used for systematic reviews (SRs) in software engineering (SE) and suggested that there should be changes to the current SR guidelines.Objective: To investigate whether current SR guidelines need to be changed to support GL and multivocal reviews.Method: We discuss the definitions of GL and the importance of GL and of industry-based field studies in SE SRs. We identify properties of SRs that constrain the material used in SRs: a) the nature of primary studies; b) the requirements of SRs to be auditable, traceable, and reproducible; and explain why these requirements restrict the use of blogs in SRs.Results: SR guidelines have always considered GL as a possible source of primary studies and have never supported exclusion of field studies that incorporate the practitioners’ viewpoint. However, the concept of GL, which was meant to refer to documents that were not formally published, is now being extended to information from sources such as blogs/tweets/Q&A posts. Thus, it might seem that SRs do not make full use of GL because they do not include such information. However, the unit of analysis for an SR is the primary study. Thus, it is not thesourcebut thetypeof information that is important. Any report describing a rigorous empirical evaluation is a candidate primary study. Whether it is actually included in an SR depends on the SR eligibility criteria. However, any study that cannot be guaranteed to be publicly available in the long term should not be used as a primary study in an SR. This does not prevent such information from being aggregated in surveys of social media and used in the context of evidence-based software engineering (EBSE).Conclusions: Current guidelines for SRs do not require extensions, but their scope needs to be better defined. SE researchers require guidelines for analysing social media posts (e.g., blogs, tweets, vlogs), but these should be based on qualitative primary (not secondary) study guidelines. SE researchers can use mixed-methods SRs and/or the fourth step of EBSE to incorporate findings from social media surveys with those from SRs and to develop industry-relevant recommendations. Barbara A. Kitchenham, Lech Madeyski, David Budgen |
IEEE Trans. Software Eng. | 1 |
| 2023 | SEGRESS: Software Engineering Guidelines for REporting Secondary StudiesabstractContext: Several tertiary studies have criticized the reporting of software engineering secondary studies.Objective: Our objective is to identify guidelines for reporting software engineering (SE) secondary studies which would address problems observed in the reporting of software engineering systematic reviews (SRs).Method: We review the criticisms of SE secondary studies and identify the major areas of concern. We assess the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) statement as a possible solution to the need for SR reporting guidelines, based on its status as the reporting guideline recommended by the Cochrane Collaboration whose SR guidelines were a major input to the guidelines developed for SE. We report its advantages and limitations in the context of SE secondary studies. We also assess reporting guidelines for mapping studies and qualitative reviews, and compare their structure and content with that of PRISMA 2020.Results: Previous tertiary studies confirm that reports of secondary studies are of variable quality. However,ad hocrecommendations that amend reporting standards may result in unnecessary duplication of text. We confirm that the PRISMA 2020 statement addresses SE reporting problems, but is mainly oriented to quantitative reviews, mixed-methods reviews and meta-analyses. However, we show that the PRISMA 2020 item definitions can be extended to cover the information needed to report mapping studies and qualitative reviews.Conclusions: In this paper and its Supplementary Material, we present and illustrate an integrated set of guidelines called SEGRESS (Software Engineering Guidelines for REporting Secondary Studies), suitable for quantitative systematic reviews (building upon PRISMA 2020), mapping studies (PRISMA-ScR), and qualitative reviews (ENTREQ and RAMESES), that addresses reporting problems found in current SE SRs. Barbara A. Kitchenham, Lech Madeyski, David Budgen |
IEEE Trans. Software Eng. | 1 |
| 2022 | The Importance of the Correlation in Crossover ExperimentsabstractContext:In empirical software engineering, crossover designs are popular for experiments comparing software engineering techniques that must be undertaken by human participants. However, their value depends on the correlation ($r$) between the outcome measures on the same participants. Software engineering theory emphasizes the importance of individual skill differences, so we would expect the values of$r$to be relatively high. However, few researchers have reported the values of$r$.Goal:To investigate the values of$r$found in software engineering experiments.Method:We undertook simulation studies to investigate the theoretical and empirical properties of$r$. Then we investigated the values of$r$observed in 35 software engineering crossover experiments.Results:The level of$r$obtained by analysing our 35 crossover experiments was small. Estimates based on means, medians, and random effect analysis disagreed but were all between 0.2 and 0.3. As expected, our analyses found large variability among the individual$r$estimates for small sample sizes, but no indication that$r$estimates were larger for the experiments with larger sample sizes that exhibited smaller variability.Conclusions:Low observed$r$values cast doubts on the validity of crossover designs for software engineering experiments. However, if the cause of low$r$values relates to training limitations or toy tasks, this affectsallSoftware Engineering (SE) experiments involving human participants. For all human-intensive SE experiments, we recommend more intensive training and then tracking the improvement of participants as they practice using specific techniques, before formally testing the effectiveness of the techniques. Barbara A. Kitchenham, Lech Madeyski, Giuseppe Scanniello, Carmine Gravino |
IEEE Trans. Software Eng. | 1 |
| 2021 | Training students in evidence-based software engineering and systematic reviews: a systematic review and empirical study
Sebastián Pizard, Fernando Acerenza, Ximena Otegui, Silvana Moreno, Diego Vallespir, Barbara A. Kitchenham |
Empir. Softw. Eng. | 6 |
| 2020 | Meta-analysis for families of experiments in software engineering: a systematic review and reproducibility and validity assessmentabstractPrevious studies have raised concerns about the analysis and meta-analysis of crossover experiments and we were aware of several families of experiments that used crossover designs and meta-analysis. To identify families of experiments that used meta-analysis, to investigate their methods for effect size construction and aggregation, and to assess the reproducibility and validity of their results. We performed a systematic review (SR) of papers reporting families of experiments in high quality software engineering journals, that attempted to apply meta-analysis. We attempted to reproduce the reported meta-analysis results using the descriptive statistics and also investigated the validity of the meta-analysis process. Out of 13 identified primary studies, we reproduced only five. Seven studies could not be reproduced. One study which was correctly analyzed could not be reproduced due to rounding errors. When we were unable to reproduce results, we provide revised meta-analysis results. To support reproducibility of analyses presented in our paper, it is complemented by the reproducer R package. Meta-analysis is not well understood by software engineering researchers. To support novice researchers, we present recommendations for reporting and meta-analyzing families of experiments and a detailed example of how to analyze a family of 4-group crossover experiments. Barbara A. Kitchenham, Lech Madeyski, Pearl Brereton |
Empir. Softw. Eng. | 1 |
| 2019 | Problems with Statistical Practice in Human-Centric Software Engineering ExperimentsabstractBackground Examples of questionable statistical practice, when published in high quality software engineering (SE) journals, may lead to novice researchers adopting incorrect statistical practices. Barbara A. Kitchenham, Lech Madeyski, Pearl Brereton |
EASE | 1 |
| 2018 | Effect sizes and their variance for AB/BA crossover design studiesabstractWe addressed the issues related to repeated measures experimental design such as an AB/BA crossover design that have been neither discussed nor addressed in the software engineering literature. Lech Madeyski, Barbara A. Kitchenham |
ICSE | 2 |
| 2018 | Effect sizes and their variance for AB/BA crossover design studiesabstractVegas et al. IEEE Trans Softw Eng 42(2):120:135 (2016) raised concerns about the use of AB/BA crossover designs in empirical software engineering studies. This paper addresses issues related to calculating standardized effect sizes and their variances that were not addressed by the Vegas et al.’s paper. In a repeated measures design such as an AB/BA crossover design each participant uses each method. There are two major implication of this that have not been discussed in the software engineering literature. Firstly, there are potentially two different standardized mean difference effect sizes that can be calculated, depending on whether the mean difference is standardized by the pooled within groups variance or the within-participants variance. Secondly, as for any estimated parameters and also for the purposes of undertaking meta-analysis, it is necessary to calculate the variance of the standardized mean difference effect sizes (which is not the same as the variance of the study). We present the model underlying the AB/BA crossover design and provide two examples to demonstrate how to construct the two standardized mean difference effect sizes and their variances, both from standard descriptive statistics and from the outputs of statistical software. Finally, we discuss the implication of these issues for reporting and planning software engineering experiments. In particular we consider how researchers should choose between a crossover design or a between groups design. Lech Madeyski, Barbara A. Kitchenham |
Empir. Softw. Eng. | 2 |
| 2018 | Introduction to the special section on Enhancing Credibility of Empirical Software Engineering
Lech Madeyski, Barbara A. Kitchenham, Krzysztof Wnuk |
Inf. Softw. Technol. | 2 |
| 2017 | Robust Statistical Methods for Empirical Software EngineeringabstractThere have been many changes in statistical theory in the past 30 years, including increased evidence that non-robust methods may fail to detect important results. The statistical advice available to software engineering researchers needs to be updated to address these issues. This paper aims both to explain the new results in the area of robust analysis methods and to provide a large-scale worked example of the new methods. We summarise the results of analyses of the Type 1 error efficiency and power of standard parametric and non-parametric statistical tests when applied to non-normal data sets. We identify parametric and non-parametric methods that are robust to non-normality. We present an analysis of a large-scale software engineering experiment to illustrate their use. We illustrate the use of kernel density plots, and parametric and non-parametric methods using four different software engineering data sets. We explain why the methods are necessary and the rationale for selecting a specific analysis. We suggest using kernel density plots rather than box plots to visualise data distributions. For parametric analysis, we recommend trimmed means, which can support reliable tests of the differences between the central location of two or more samples. When the distribution of the data differs among groups, or we have ordinal scale data, we recommend non-parametric methods such as Cliff’s δ or a robust rank-based ANOVA-like method. Barbara A. Kitchenham, Lech Madeyski, David Budgen, Jacky W. Keung, Pearl Brereton, Stuart M. Charters, Shirley Gibbs, Amnart Pohthong |
Empir. Softw. Eng. | 1 |
| 2015 | Robust statistical methods: why, what and how: keynoteabstractThis keynote discusses the need for more robust statistical methods. For visualizing data I suggest using Kernel density plots rather than box plots. For parametric analysis, I propose more robust measures of central location such as trimmed means, which can support reliable tests of the differences between the central location of two or more samples. In addition, I also recommend non-parametric effect sizes such as Cliff's δ and Brunner and Munzel's p-hat that avoid some of the problems with rank-based non-parametric methods. Barbara A. Kitchenham |
EASE | 1 |
| 2015 | Tools to support systematic reviews in software engineering: a cross-domain survey using semi-structured interviewsabstractBackground: A number of software tools are being developed to support systematic reviewers within the software engineering domain. However, at present, we are not sure which aspects of the review process can most usefully be supported by such tools or what characteristics of the tools are most important to reviewers. Aim: The aim of the study is to explore the scope and practice of tool support for systematic reviewers in other disciplines. Method: Researchers with experience of performing systematic reviews in Healthcare and the Social Sciences were surveyed. Qualitative data was collected through semi-structured interviews and data analysis followed an inductive approach. Results: 13 interviews were carried out. 21 software tools categorised into one of seven types were identified. Reference managers were the most commonly mentioned tools. Features considered particularly important by participants were support for multiple users, support for data extraction and support for tool maintenance. The features and importance levels identified by participants were compared with those proposed for tools to support systematic reviews in software engineering. Conclusions: Many problems faced by systematic reviewers in other disciplines are similar to those faced in software engineering. There is general consensus across domains that improved tools are needed. Christopher Marshall, Pearl Brereton, Barbara A. Kitchenham |
EASE | 3 |
| 2015 | Quality of service approaches in cloud computing: A systematic mapping study
Abdelzahir Abdelmaboud, Dayang N. A. Jawawi, Imran Ghani, Abubakar Elsafi, Barbara A. Kitchenham |
J. Syst. Softw. | 5 |
| 2014 | Tools to support systematic reviews in software engineering: a feature analysisabstractBackground The labour intensive and error prone nature of the systematic review process has led to the development and use of a range of tools to provide automated support. Christopher Marshall, Pearl Brereton, Barbara A. Kitchenham |
EASE | 3 |
| 2014 | Risks and risk mitigation in global software development: A tertiary study
June M. Verner, Pearl Brereton, Barbara A. Kitchenham, Mahmood Turner, Mahmood Niazi |
Inf. Softw. Technol. | 3 |
| 2013 | The Case for Knowledge TranslationabstractContext: For the outcomes of systematic literature reviews to be of use for practitioners, we need to develop models for addressing the needs of Knowledge Translation (KT). Aim: To identify some of the key issues that need to be addressed by a KT process for software engineering (SE) and possible routes for achieving these. Method: We have examined some of the models used in other disciplines, and suggested a possible interpretation for software engineering. Results: We propose a model for achieving KT. Conclusions: Research with industry and commerce is needed to explore how this can be realised. David Budgen, Barbara A. Kitchenham, Pearl Brereton |
ESEM | 2 |
| 2013 | Lessons from Conducting a Distributed Quasi-experimentabstractContext: Due to the lack of suitably skilled participants, software engineering experiments often lack the statistical power needed to detect the levels of effect that may be encountered. Aim: To investigate whether this can be remedied by running an experiment across multiple sites, organised as a single study rather than as a set of replications. Method: We performed a `trial' of the idea using a topic (structured abstracts) that some of us had studied previously and which required no participant training. We used five sites, each with 16 participants. Results: We were able to demonstrate the benefits of increased statistical power (and of structured abstracts). We report on our experiences with designing and conducting the study and identify some key lessons about how future studies of this form might be organised. Conclusions: The distributed model offers a flexible, robust form that is capable of delivering better statistical power than would be achieved by running a set of parallel replicated studies. David Budgen, Barbara A. Kitchenham, Stuart M. Charters, Shirley Gibbs, Amnart Pohthong, Jacky W. Keung, Pearl Brereton |
ESEM | 2 |
| 2013 | A systematic review of systematic review process research in software engineering
Barbara A. Kitchenham, Pearl Brereton |
Inf. Softw. Technol. | 1 |
| 2013 | Trends in the Quality of Human-Centric Software Engineering Experiments-A Quasi-ExperimentabstractContext: Several text books and papers published between 2000 and 2002 have attempted to introduce experimental design and statistical methods to software engineers undertaking empirical studies. Objective: This paper investigates whether there has been an increase in the quality of human-centric experimental and quasi-experimental journal papers over the time period 1993 to 2010. Method: Seventy experimental and quasi-experimental papers published in four general software engineering journals in the years 1992-2002 and 2006-2010 were each assessed for quality by three empirical software engineering researchers using two quality assessment methods (a questionnaire-based method and a subjective overall assessment). Regression analysis was used to assess the relationship between paper quality and the year of publication, publication date group (before 2003 and after 2005), source journal, average coauthor experience, citation of statistical text books and papers, and paper length. The results were validated both by removing papers for which the quality score appeared unreliable and using an alternative quality measure. Results: Paper quality was significantly associated with year, citing general statistical texts, and paper length (p <; 0.05). Paper length did not reach significance when quality was measured using an overall subjective assessment. Conclusions: The quality of experimental and quasi-experimental software engineering papers appears to have improved gradually since 1993. Barbara A. Kitchenham, Dag I. K. Sjøberg, Tore Dybå, Pearl Brereton, David Budgen, Martin Höst, Per Runeson |
IEEE Trans. Software Eng. | 1 |
| 2012 | Mapping study completeness and reliability - a case studyabstractContext: We have been undertaking a series of case studies to investigate the value of mapping (scoping) studies in software engineering. Our previous studies have assessed these using the subjective opinions of researchers. Objective: In order to provide a more objective assessment of value, for this study, we used the results of a systematic mapping study to investigate how well mapping studies identify clusters of related studies and to what extent such clusters are complete. Method: In this participant-observer case study, we undertook a mapping study of unit testing and regression testing empirical studies, which we compared with a previous expert literature review and with six other mapping studies and systematic literature reviews (SLRs) that addressed overlapping topics. Results: Our mapping study found more clusters than the expert literature review although it benefited from the set of studies identified by the expert review when refining our search process. The set of studies found by our searches were less complete than those found by SLRs addressing more specific topics, although we found some studies missed by those SLRs. Conclusions: Researchers undertaking systematic reviews and mapping studies should make use of related systematic reviews and mapping studies to identify known studies in order to refine search strings and validate search results. For completeness and traceability, mapping studies should keep a record of all multiple reports of a single study. Meta-analyses and other systematic literature reviews undertaking detailed aggregation should report on candidate primary studies that were rejected in the final screening process, as well as candidate studies that were included. This helps ensure the repeatability of aggregation results. Barbara A. Kitchenham, Pearl Brereton, David Budgen |
EASE | 1 |
| 2012 | Systematic literature reviews in global software development: A tertiary studyabstractContext: There has been an increase in research into global software development (GSD) and in the number of systematic literature reviews (SLRs) addressing this topic. Objective: The aim of this research is to catalogue GSD SLRs in order to identify the topics covered, the active researchers, the publication vehicles, and to assess the quality of the SLRs identified. Method: We performed a broad automated search to find SLRs dealing with GSD. We differentiate between SLR studies and papers reporting those studies. Data relating to each of the following was extracted and synthesized from each study: authors and their affiliation at the time of publication, the journal or conference in which the paper was published, the quality of each study and the main GSD study topic. Results: Twenty-four GSD SLR studies and 37 papers reporting those studies were identified. Major GSD topics covered include: (1) organizational environment, (2) project execution, and (3) project planning and control. The main research groups are based in Brazil (17), Ireland (8), and Sweden (7). Conclusions: GSD SLR studies are most frequently reported in the International Conference on Global Software Engineering and IEEE Software; the two most popular topics for research are risk factors due to the organizational environment and the development process. The most active researchers are based in Brazil. The quality of the SLR studies has not changed over time. June M. Verner, Pearl Brereton, Barbara A. Kitchenham, Mark Turner 0001, Mahmood Niazi |
EASE | 3 |
| 2012 | Three empirical studies on the agreement of reviewers about the quality of software engineering experimentsabstractDuring systematic literature reviews it is necessary to assess the quality of empirical papers. Current guidelines suggest that two researchers should independently apply a quality checklist and any disagreements must be resolved. However, there is little empirical evidence concerning the effectiveness of these guidelines. This paper investigates the three techniques that can be used to improve the reliability (i.e. the consensus among reviewers) of quality assessments, specifically, the number of reviewers, the use of a set of evaluation criteria and consultation among reviewers. We undertook a series of studies to investigate these factors. Two studies involved four research papers and eight reviewers using a quality checklist with nine questions. The first study was based on individual assessments, the second study on joint assessments with a period of inter-rater discussion. A third more formal randomised block experiment involved 48 reviewers assessing two of the papers used previously in teams of one, two and three persons to assess the impact of discussion among teams of different size using the evaluations of the “teams” of one person as a control. For the first two studies, the inter-rater reliability was poor for individual assessments, but better for joint evaluations. However, the results of the third study contradicted the results of Study 2. Inter-rater reliability was poor for all groups but worse for teams of two or three than for individuals. When performing quality assessments for systematic literature reviews, we recommend using three independent reviewers and adopting the median assessment. A quality checklist seems useful but it is difficult to ensure that the checklist is both appropriate and understood by reviewers. Furthermore, future experiments should ensure participants are given more time to understand the quality checklist and to evaluate the research papers. Barbara A. Kitchenham, Dag I. K. Sjøberg, Tore Dybå, Dietmar Pfahl, Pearl Brereton, David Budgen, Martin Höst, Per Runeson |
Inf. Softw. Technol. | 1 |
| 2012 | Interpretation problems related to the use of regression models to decide on economy of scale in software development
Magne Jørgensen, Barbara A. Kitchenham |
J. Syst. Softw. | 2 |
| 2012 | Toward trustworthy software process models: an exploratory study on transformable process modelingabstractSUMMARY Software process modeling and simulation have become effective tools for support of software process management and improvement over the past two decades. They have recently been integrated into the Trustworthy Process Management Framework (TPMF) as the infrastructural components to facilitate the delivery of trustworthy software products. This paper proposes the concept of Trustworthy Software Process Models as inputs to TPMF and introduces transformable process modeling for supporting effective and productive development of trustworthy process models. Furthermore, this paper undertakes an exploratory study on process model transformation by investigating and comparing process modeling semantics between quantitative (e.g., System Dynamics, SD) and qualitative forms of modeling and simulation. By following the model transformation scheme, a quantitative continuous (SD) software evolution process model is successfully transformed into its qualitative form for simulation. The results present the different capabilities and performance between these two modeling paradigms, as well as the possible benefits and interesting perspectives of transformable process modeling. Copyright © 2010 John Wiley & Sons, Ltd. He Zhang 0001, Barbara A. Kitchenham, D. Ross Jeffery |
J. Softw. Evol. Process. | 2 |
| 2011 | Repeatability of systematic literature reviewsabstractBackground: One of the anticipated benefits of systematic literature reviews (SLRs) is that they can be conducted in an auditable way to produce repeatable results. Aim: This study aims to identify under what conditions SLRs are likely to be stable, with respect to the primary studies selected, when used in software engineering. The conditions we investigate in this report are when novice researchers undertake searches with a common goal. Method: We undertook a participant-observer multi-case study to investigate the repeatability of systematic literature reviews. The 'cases' in this study were the early stages, involving identification of relevant literature, of two SLRs of unit testing methods. The SLRs were performed independently by two novice researchers. The SLRs were restricted to the ACM and IEEE digital libraries for the years 1986-2005 so their results could be compared with a published expert literature review of unit testing papers. Results: The two SLRs selected very different papers with only six papers out of 32 in common, and both differed substantially from a published secondary study of unit testing papers finding only three of 21 papers. Of the 29 additional papers found by the novice researchers, only 10 were considered relevant. The 10 additional relevant papers would have had an impact on the results of the published study by adding three new categories to the framework and adding papers to three, otherwise empty, cells. Conclusions: In the case of novice researchers, having broadly the same research question will not necessarily guarantee repeatability with respect to primary studies. Systematic reviews must be careful to report their search process fully or they will not be repeatable. Missing papers can have a significant impact on the stability of the results of a secondary study. Barbara A. Kitchenham, Pearl Brereton, Zhi Li 0017, David Budgen, Andrew James Burn |
EASE | 1 |
| 2011 | Reporting computing projects through structured abstracts: a quasi-experiment
David Budgen, Andy J. Burn, Barbara A. Kitchenham |
Empir. Softw. Eng. | 3 |
| 2011 | Using mapping studies as the basis for further research - A participant-observer case study
Barbara A. Kitchenham, David Budgen, Pearl Brereton |
Inf. Softw. Technol. | 1 |
| 2011 | Empirical evidence about the UML: a systematic literature reviewabstractAbstract The Unified Modeling Language (UML) was created on the basis of expert opinion and has now become accepted as the ‘standard’ object‐oriented modelling notation. Our objectives were to determine how widely the notations of the UML, and their usefulness, have been studied empirically, and to identify which aspects of it have been studied in most detail. We undertook a mapping study of the literature to identify relevant empirical studies and to classify them in terms of the aspects of the UML that they studied. We then conducted a systematic literature review, covering empirical studies published up to the end of 2008, based on the main categories identified. We identified 49 relevant publications, and report the aggregated results for those categories for which we had enough papers—metrics,comprehension,model quality,methods and toolsandadoption. Despite indications that a number of problems exist with UML models, researchers tend to use the UML as a ‘given’ and seem reluctant to ask questions that might help to make it more effective. Copyright © 2010 John Wiley & Sons, Ltd. David Budgen, Andy J. Burn, Pearl Brereton, Barbara A. Kitchenham, Rialette Pretorius |
Softw. Pract. Exp. | 4 |
| 2010 | The value of mapping studies - A participant-observer case study
Barbara A. Kitchenham, David Budgen, Pearl Brereton |
EASE | 1 |
| 2010 | Can we evaluate the quality of software engineering experiments?abstractContext: The authors wanted to assess whether the quality of published human-centric software engineering experiments was improving. This required a reliable means of assessing the quality of such experiments. Aims: The aims of the study were to confirm the usability of a quality evaluation checklist, determine how many reviewers were needed per paper that reports an experiment, and specify an appropriate process for evaluating quality. Method: With eight reviewers and four papers describing human-centric software engineering experiments, we used a quality checklist with nine questions. We conducted the study in two parts: the first was based on individual assessments and the second on collaborative evaluations. Results: The inter-rater reliability was poor for individual assessments but much better for joint evaluations. Four reviewers working in two pairs with discussion were more reliable than eight reviewers with no discussion. The sum of the nine criteria was more reliable than individual questions or a simple overall assessment. Conclusions: If quality evaluation is critical, more than two reviewers are required and a round of discussion is necessary. We advise using quality criteria and basing the final assessment on the sum of the aggregated criteria. The restricted number of papers used and the relatively extensive expertise of the reviewers limit our results. In addition, the results of the second part of the study could have been affected by removing a time restriction on the review as well as the consultation process. Barbara A. Kitchenham, Dag I. K. Sjøberg, Pearl Brereton, David Budgen, Tore Dybå, Martin Höst, Dietmar Pfahl, Per Runeson |
ESEM | 1 |
| 2010 | The educational value of mapping studies of software engineering literatureabstractWe identify three challenges related to the provenance of the material we use in teaching software engineering. We suggest that these challenges can be addressed by using evidence-based software engineering (EBSE) and its primary tool of systematic literature reviews (SLRs). This paper aims to assess the educational and scientific value of undergraduate and postgraduate students undertaking a specific form of SLR called a mapping study. Using a case study methodology, we asked three postgraduate students and three undergraduates and their supervisor to complete a questionnaire concerning the educational value of mapping studies and any problems they experienced. Students found undertaking a mapping study to be a valuable experience providing both reusable research skills and a good overview of a research topic. Postgraduates found it useful as a starting point for their studies. Undergraduates reported problems undertaking the study in the required timescales. Searching and classifying the literature was difficult. Barbara A. Kitchenham, Pearl Brereton, David Budgen |
ICSE (1) | 1 |
| 2010 | Refining the systematic literature review process - two participant-observer case studies
Barbara A. Kitchenham, Pearl Brereton, Mark Turner 0001, Mahmood Niazi, Stephen G. Linkman, Rialette Pretorius, David Budgen |
Empir. Softw. Eng. | 1 |
| 2010 | Evaluating logistic regression models to estimate software project outcomes
Narciso Cerpa, Matthew Bardeen, Barbara A. Kitchenham, June M. Verner |
Inf. Softw. Technol. | 3 |
| 2010 | Systematic literature reviews in software engineering - A tertiary study
Barbara A. Kitchenham, Rialette Pretorius, David Budgen, Pearl Brereton, Mark Turner 0001, Mahmood Niazi, Stephen G. Linkman |
Inf. Softw. Technol. | 1 |
| 2010 | Does the technology acceptance model predict actual use? A systematic literature review
Mark Turner 0001, Barbara A. Kitchenham, Pearl Brereton, Stuart M. Charters, David Budgen |
Inf. Softw. Technol. | 2 |
| 2010 | What's up with software metrics? - A preliminary mapping study
Barbara A. Kitchenham |
J. Syst. Softw. | 1 |
| 2010 | How Reliable Are Systematic Reviews in Empirical Software Engineering?abstractBACKGROUND-The systematic review is becoming a more commonly employed research instrument in empirical software engineering. Before undue reliance is placed on the outcomes of such reviews it would seem useful to consider the robustness of the approach in this particular research context. OBJECTIVE-The aim of this study is to assess the reliability of systematic reviews as a research instrument. In particular, we wish to investigate the consistency of process and the stability of outcomes. METHOD-We compare the results of two independent reviews undertaken with a common research question. RESULTS-The two reviews find similar answers to the research question, although the means of arriving at those answers vary. CONCLUSIONS-In addressing a well-bounded research question, groups of researchers with similar domain experience can arrive at the same review outcomes, even though they may do so in different ways. This provides evidence that, in this context at least, the systematic review is a robust research method. Stephen G. MacDonell, Martin J. Shepperd, Barbara A. Kitchenham, Emilia Mendes |
IEEE Trans. Software Eng. | 3 |
| 2009 | A Quality Checklist for Technology-Centred Testing Studies
Barbara A. Kitchenham, Andrew James Burn, Zhi Li 0017 |
EASE | 1 |
| 2009 | An Evaluation of Quality Checklist Proposals - A participant-observer case study
Barbara A. Kitchenham, Pearl Brereton, David Budgen, Zhi Li 0017 |
EASE | 1 |
| 2009 | The impact of limited search procedures for systematic literature reviews A participant-observer case studyabstractThis study aims to compare the use of targeted manual searches with broad automated searches, and to assess the importance of grey literature and breadth of search on the outcomes of SLRs. We used a participant-observer multi-case embedded case study. Our two cases were a tertiary study of systematic literature reviews published between January 2004 and June 2007 based on a manual search of selected journals and conferences and a replication of that study based on a broad automated search. Broad searches find more papers than restricted searches, but the papers may be of poor quality. Researchers undertaking SLRs may be justified in using targeted manual searches if they intend to omit low quality papers; if publication bias is not an issue; or if they are assessing research trends in research methodologies. Barbara A. Kitchenham, Pearl Brereton, Mark Turner 0001, Mahmood Niazi, Stephen G. Linkman, Rialette Pretorius, David Budgen |
ESEM | 1 |
| 2009 | Guidelines for Industrially-Based Multiple Case Studies in Software EngineeringabstractWithout careful methodological guidance, case studies in software engineering are difficult to plan, design and execute. While there are a number of broad guidelines for case study research, there are none that specifically address the needs of a software engineer undertaking multiple case studies in an industrial setting. Through a synthesis of existing best practices in case study research, we provide a set of comprehensive guidelines for conducting multiple case studies in software engineering research. Our guidelines can assist software engineering researchers with all stages of multiple case study research, although in this paper we concentrate on the early phases, such as focusing the case study and detailed plan design. To date, three exploratory research projects found our guidelines very useful. We illustrate our guidelines with examples from one of these projects. June M. Verner, Jennifer Sampson, Vladimir Tosic, Nur Azzah Abu Bakar, Barbara A. Kitchenham |
RCIS | 5 |
| 2009 | Systematic literature reviews in software engineering - A systematic literature review
Barbara A. Kitchenham, Pearl Brereton, David Budgen, Mark Turner 0001, John Bailey, Stephen G. Linkman |
Inf. Softw. Technol. | 1 |
| 2008 | Software Process Simulation Modeling: Facts, Trends and DirectionsabstractSoftware process simulation modeling (SPSM) research has increased since the first ProSim workshop held in 1998 and Kellner, Madachy and Raffo (KMR) discussed the "why, what and how" of process simulation. This paper aims to assess how SPSM has evolved during the past 10 years in particular whether the reasons for SPSM, the simulation paradigms, tools, problem domains, and model scopes have changed. We performed a systematic literature review of software process simulation papers from the ProSim series publications in the last decade. We identified 96 studies from the sources and included them in this review. The papers were categorized into four major types and data needed to address each research question was extracted. We found a need for refining the reasons and the classification scheme for SPSM introduced by KMR. More emerging SPSM paradigms and model scopes were added to enhance KMR's discussion. Trends over time showed that interest in continuous modeling was decreasing and interest in micro-processes was increasing. Hybrid models were based primarily on system dynamics and discrete event simulation and were all implemented by vertical integration. We recommend SPSM research concentrate more on recent software processes and on making SPSM more reusable and thus easier to build. He Zhang 0001, Barbara A. Kitchenham, Dietmar Pfahl |
APSEC | 2 |
| 2008 | Using a Protocol Template for Case Study Planning
Pearl Brereton, Barbara A. Kitchenham, David Budgen, Zhi Li 0017 |
EASE | 2 |
| 2008 | Lessons from a cross domain investigation of empirical practices
David Budgen, John Bailey, Mark Turner 0001, Barbara A. Kitchenham, Pearl Brereton, Stuart M. Charters |
EASE | 4 |
| 2008 | Lessons learnt Undertaking a Large-scale Systematic Literature Review
Mark Turner 0001, Barbara A. Kitchenham, David Budgen, Pearl Brereton |
EASE | 2 |
| 2008 | Software process simulation over the past decade: trends discovery from a systematic reviewabstractSoftware Process Simulation (SPS) research has increased since 1998 when the first ProSim Workshop was held. This paper aims to reveal how SPS has evolved during the past 10 years based on the preliminary results from the systematic literature review of SPS publications from 1998 to 2007. Trends over the period showed that interest in continuous modelling was decreasing and interest in micro-processes was increasing. Hybrid models were based primarily on system dynamics and discrete event simulation and were all implemented by vertical integration. He Zhang 0001, Barbara A. Kitchenham, Dietmar Pfahl |
ESEM | 2 |
| 2008 | Comparing distributed and face-to-face meetings for software architecture evaluation: A controlled experiment
Muhammad Ali Babar 0001, Barbara A. Kitchenham, D. Ross Jeffery |
Empir. Softw. Eng. | 2 |
| 2008 | Presenting software engineering results using structured abstracts: a randomised experiment
David Budgen, Barbara A. Kitchenham, Stuart M. Charters, Mark Turner 0001, Pearl Brereton, Stephen G. Linkman |
Empir. Softw. Eng. | 2 |
| 2008 | The role of replications in empirical software engineering - a word of warning
Barbara A. Kitchenham |
Empir. Softw. Eng. | 1 |
| 2008 | Evaluating guidelines for reporting empirical software engineering studies
Barbara A. Kitchenham, Hiyam Al-Kilidar, Muhammad Ali Babar 0001, Mike Berry, Karl Cox, Jacky W. Keung, Felicia Kurniawati, Mark Staples, He Zhang 0001, Liming Zhu 0001 |
Empir. Softw. Eng. | 1 |
| 2008 | Analogy-X: Providing Statistical Inference to Analogy-Based Software Cost EstimationabstractAbstract Data-intensive analogy has been proposed as a means of software cost estimation as an alternative to other data intensive methods such as linear regression. Unfortunately, there are drawbacks to the method. There is no mechanism to assess its appropriateness for a specific dataset. In addition, heuristic algorithms are necessary to select the best set of variables and identify abnormal project cases. We introduce a solution to these problems based upon the use of the Mantel correlation randomization test called Analogy-X. We use the strength of correlation between the distance matrix of project features and the distance matrix of known effort values of the dataset. The method is demonstrated using the Desharnais dataset and two random datasets, showing (1) the use of Mantel's correlation to identify whether analogy is appropriate, (2) a stepwise procedure for feature selection, as well as (3) the use of a leverage statistic for sensitivity analysis that detects abnormal data points. Analogy-X, thus, provides a sound statistical basis for analogy, removes the need for heuristic search and greatly improves its algorithmic performance. Jacky W. Keung, Barbara A. Kitchenham, D. Ross Jeffery |
IEEE Trans. Software Eng. | 2 |
| 2007 | Optimising Project Feature Weights for Analogy-Based Software Cost Estimation using the Mantel CorrelationabstractSoftware cost estimation using analogy is an important area in software engineering research. Previous research has demonstrated that analogy is a viable alternative to other conventional estimation methods in terms of predictive accuracy. One of the important research areas for analogy is how to determine suitable project feature weights. This can be achieved by using an extensive project feature weights search, where the quality measure is optimised. However, this approach suffers similar issues as the brute-force feature selection approach in analogy. We propose a novel method to deal with this issue based upon the use of the Mantel randomisation test. Specifically, we determine project feature weights based on the strength of correlation between the distance matrix of project features and the distance matrix of known effort values of the dataset. We demonstrate the procedure on a specific dataset, showing the use of the Mantel correlation to identify whether analogy is appropriate, and whether the project feature weights can be determined by statistical inference. Our results also show improved prediction accuracy when multiple project features are used with determined weights. Our method, thus, provides a sound statistical basis for analogy. Jacky W. Keung, Barbara A. Kitchenham |
APSEC | 2 |
| 2007 | Assessment of a Framework for Comparing Software Architecture Analysis Methods
Muhammad Ali Babar 0001, Barbara A. Kitchenham |
EASE | 2 |
| 2007 | Systematic Review of Statistical Process Control: An Experience Report
Maria Teresa Baldassarre, Danilo Caivano, Barbara A. Kitchenham, Giuseppe Visaggio |
EASE | 3 |
| 2007 | Preliminary results of a study of the completeness and clarity of structured abstracts
David Budgen, Barbara A. Kitchenham, Stuart M. Charters, Mark Turner 0001, Pearl Brereton, Stephen G. Linkman |
EASE | 2 |
| 2007 | The Impact of Group Size on Software Architecture Evaluation: A Controlled ExperimentabstractAn important element in scenario-based architecture evaluation is the development of scenario profiles by stakeholders working in groups. In practice groups can vary in size from 2 to 20 people. Currently, there is no empirical evidence about the impact of group size on the scenario development activity. Our experimental goal was to investigate the impact of group size on the quality of scenario profiles developed by different sizes of groups. We had 165 subjects, who were randomly assigned to 10 groups of size 3, 13 groups of size 5, and 10 groups of size 7. Participants were asked to develop scenario profiles. After the experiment each participant completed a questionnaire aimed at identifying their opinion of the group activity. The average quality score for group scenario profiles for 3 person groups was 362.4, for groups of 5 person groups was 534.23 and for 7 person groups was. 444.5. The quality of scenario profiles for groups of size 5 was significantly greater than the quality of scenario profiles for groups of size 3 (p=0.025), but there was no difference between the size 3 and size 7 groups. However, participants in groups of size 3 had a significantly better opinion of the group activity outcome and their personal interaction with their group than participants in groups of size 5 or 7. Our results suggest that the quality of the output from a group does not increase linearly with group size. However, individual participants prefer small groups. This means there is a trade-off between group output quality and the personal experience of group members. Muhammad Ali Babar 0001, Barbara A. Kitchenham |
ESEM | 2 |
| 2007 | Evidence relating to Object-Oriented software design: A surveyabstractThere is little empirical knowledge of the effectiveness of the object-oriented paradigm. To conduct a systematic review of the literature describing empirical studies of this paradigm. We undertook a Mapping Study of the literature. 138 papers have been identified and classified by topic, form of study involved, and source. The majority of empirical studies of OO (object oriented software) concentrate on metrics, relatively few consider effectiveness. John Bailey, David Budgen, Mark Turner 0001, Barbara A. Kitchenham, Pearl Brereton, Stephen G. Linkman |
ESEM | 4 |
| 2007 | Lessons from applying the systematic literature review process within the software engineering domainabstractA consequence of the growing number of empirical studies in software engineering is the need to adopt systematic approaches to assessing and aggregating research outcomes in order to provide a balanced and objective summary of research evidence for a particular topic. The paper reports experiences with applying one such approach, the practice of systematic literature review, to the published studies relevant to topics within the software engineering domain. The systematic literature review process is summarised, a number of reviews being undertaken by the authors and others are described and some lessons about the applicability of this practice to software engineering are extracted. The basic systematic literature review process seems appropriate to software engineering and the preparation and validation of a review protocol in advance of a review activity is especially valuable. The paper highlights areas where some adaptation of the process to accommodate the domain-specific characteristics of software engineering is needed as well as areas where improvements to current software engineering infrastructure and practices would enhance its applicability. In particular, infrastructure support provided by software engineering indexing databases is inadequate. Also, the quality of abstracts is poor; it is usually not possible to judge the relevance of a study from a review of the abstract alone. Pearl Brereton, Barbara A. Kitchenham, David Budgen, Mark Turner 0001, Mohamed Khalil |
J. Syst. Softw. | 2 |
| 2007 | Introduction to special section on Evaluation and Assessment in Software Engineering EASE06
Barbara A. Kitchenham, Pearl Brereton |
J. Syst. Softw. | 1 |
| 2007 | Cross versus Within-Company Cost Estimation Studies: A Systematic ReviewabstractThe objective of this paper is to determine under what circumstances individual organizations would be able to rely on cross-company-based estimation models. We performed a systematic review of studies that compared predictions from cross-company models with predictions from within-company models based on analysis of project data. Ten papers compared cross-company and within-company estimation models; however, only seven presented independent results. Of those seven, three found that cross-company models were not significantly different from within-company models, and four found that cross-company models were significantly worse than within-company models. Experimental procedures used by the studies differed making it impossible to undertake formal meta-analysis of the results. The main trend distinguishing study results was that studies with small within-company data sets (i.e., $20 projects) that used leave-one-out cross validation all found that the within-company model was significantly different (better) from the cross-company model. The results of this review are inconclusive. It is clear that some organizations would be ill-served by cross-company models whereas others would benefit. Further studies are needed, but they must be independent (i.e., based on different data bases or at least different single company data sets) and should address specific hypotheses concerning the conditions that would favor cross-company or within-company models. In addition, experimenters need to standardize their experimental procedures to enable formal meta-analysis, and recommendations are made in Section 3. Barbara A. Kitchenham, Emilia Mendes, Guilherme Horta Travassos |
IEEE Trans. Software Eng. | 1 |
| 2006 | Assessing the Value of Architectural Information Extracted from Patterns for ArchitectingabstractBackground: We have developed an approach to identifying and capturing architecturally significant information from patterns (ASIP), which can be used to improve architecture design and evaluation. Goal: Our goal was to evaluate whether the use of the ASIP provides more effective support in understanding or designing software architecture composed of the software design patterns which are the source of the ASIP compared with the original design pattern documentation. Experimental design: Our subjects were 20 experienced software engineers who had returned to University for a post graduate course. All participants were taking a course in software architecture. The participants were randomly assigned to two groups of equal size. Both groups performed two tasks: understanding the use of J2EE design pattern in a given architecture based on the quality requirements the architecture was supported to satisfy, and designing software architecture to satisfy a given set of quality requirements using J2EE design patterns. For the first task, one group (treatment group) was given ASIP information the other (control group) was given the standard J2EE pattern documentation. For the second task, treatment group became the control group and vice versa and the type of support information was kept constant. The outcome variables were the number of correctly identified design patterns. The participants also completed a post-experiment questionnaire. Result: The average score for the first task for the treatment group was 23.90 and for the control group was 13.80. The difference between the groups was significant using Mann-Whiney test (p=0.0375). The average score for the second task for the treatment group was 26.85 and for the control group was 19.60. Mann-Whitney test revealed that the difference between the groups was again significant at (p=0.035). Post-study questionnaire revealed that 18 of the 20 participants believed that the ASIP was more helpful than pattern documentation for understanding and designing architectures. Conclusion: Our results support the hypotheses that ASIP information is more helpful in understanding or designing software architectures using software design patterns than pattern documentation itself. Muhammad Ali Babar 0001, Barbara A. Kitchenham |
EASE | 2 |
| 2006 | A Systematic review of Cross- vs. Within-Company Cost Estimation StudiesabstractOBJECTIVE – The objective of this paper is to determine under what circumstances individual organisations would be able to rely on cross-company based estimation models. METHOD – We performed a systematic review of studies that compared predictions from cross-company models with predictions from within-company models based on analysis of project data. RESULTS – Ten papers compared cross-company and within-company estimation models, however, only seven of the papers presented independent results. Of those seven, three found that cross-company models were as good as within-company models, four found cross-company models were significantly worse than within-company models. Experimental procedures used by the studies differed making it impossible to undertake formal meta-analysis of the results. The main trend distinguishing study results was that studies with small single company data sets (i.e. <20 projects) that used leave-one-out cross-validation all found that the within-company model was significantly more accurate than the cross-company model. CONCLUSIONS – The results of this review are inconclusive. It is clear that some organisations would be ill-served by cross-company models whereas others would benefit. Further studies are needed, but they must be independent (i.e. based on different data bases or at least different single company data sets). In addition, experimenters need to standardise their experimental procedures to enable formal meta-analysis. Barbara A. Kitchenham, Emilia Mendes, Guilherme Horta Travassos |
EASE | 1 |
| 2006 | Towards a distributed software architecture evaluation process: a preliminary assessmentabstractScenario-based methods for evaluating software architecture require a large number of stakeholders to be collocated for evaluation sessions. Collocating stakeholders is often an expensive exercise. We have proposed a framework for distributed evaluation process. We present the proposed framework and initial results of a controlled experiment that we ran to assess the effectiveness of the proposed idea. Muhammad Ali Babar 0001, Barbara A. Kitchenham, Ian Gorton |
ICSE | 2 |
| 2006 | Lessons learnt from the analysis of large-scale corporate databasesabstractThis paper presents the lessons learnt during the analysis of the corporate databases developed by IBM Global Services (Australia). IBM is rated as CMM level 5. Following CMM level 4 and above practices, IBM designed several software metrics databases with associated data collection and reporting systems to manage its corporate goals. However, IBM quality staff believed the data were not as useful as they had expected. NICTA staff undertook a review of IBM's statistical process control procedures and found problems with the databases mainly due to a lack of links between the different data tables. Such problems might be avoided by using M3P variant of the GQM paradigm to define a hierarchy of goals, with project goals at the lowest level, then process goals and corporate goals at the highest level. We propose using E-R models to identify problems with existing databases and to design databases once goals have been defined. Barbara A. Kitchenham, Cat Kutay, D. Ross Jeffery, Colin Connaughton |
ICSE | 1 |
| 2006 | Evidence-Based Software Engineering and Systematic Literature Reviews
Barbara A. Kitchenham |
PROFES | 1 |
| 2006 | Evaluation and Assessment in Software Engineering (EASE 05)
Barbara A. Kitchenham |
Inf. Softw. Technol. | 1 |
| 2006 | An empirical study of groupware support for distributed software architecture evaluation processabstractSoftware architecture evaluation is an effective means of addressing quality related issues early in the software development lifecycle. Scenario-based approaches to evaluate architecture usually involve a large number of stakeholders, who need to be collocated for face-to-face evaluation meetings. Collocating a large number of stakeholders is an expensive and time-consuming exercise, which may prove to be a hurdle in the wide-spread adoption of disciplined architectural evaluation practices. Drawing upon the successful introduction of groupware applications to support geographically distributed teams in software inspection, and requirements engineering disciplines, we propose the concept of distributed architectural evaluation using Internet-based collaborative technologies. This paper presents a pilot study used to assess the viability of a larger experiment intended to investigate the feasibility of groupware support for distributed software architecture evaluation. In addition, the results of the pilot study provide some preliminary findings on the viability of groupware-supported software architectural evaluation process. Muhammad Ali Babar 0001, Barbara A. Kitchenham, Liming Zhu 0001, Ian Gorton, D. Ross Jeffery |
J. Syst. Softw. | 2 |
| 2005 | International workshop on realising evidence-based software engineeringabstractThis workshop is concerned with defining the procedures that are needed to establish a sound empirical foundation for the practices of Software Engineering. Our goal is to begin building a community that will review, analyse, codify and promulgate software engineering experiences as well as to identify the processes and infrastructure that are needed to support these activities. David Budgen, Pearl Brereton, Barbara A. Kitchenham, Stephen G. Linkman |
ICSE | 3 |
| 2005 | Experiences of using an evaluation frameworkabstractThis paper reports two trials of an evaluation framework intended to evaluate novel software applications. The evaluation framework was originally developed to evaluate a risk-based software bidding model, and our first trial of using the framework was our evaluation of the bidding model. We found that the framework worked well as a validation framework but needed to be extended before it would be appropriate for evaluation. Subsequently, we compared our framework with a recently completed evaluation of a software tool undertaken as part of the Framework V CLARiFi project. In this case, we did not use the framework to guide the evaluation; we used the framework to see whether it would identify any weaknesses in the actual evaluation process. Activities recommended by the framework were not undertaken in the order suggested by the evaluation process and we found problems relating to that oversight surfaced during the tool evaluation activities. Our experiences suggest that the framework has some benefits but it also requires further practical testing. Barbara A. Kitchenham, Stephen G. Linkman, Susan Linkman |
Inf. Softw. Technol. | 1 |
| 2005 | A framework for evaluating a software bidding modelabstractThis paper discusses the issues involved in evaluating a software bidding model. We found it difficult to assess the appropriateness of any model evaluation activities without a baseline or standard against which to assess them. This paper describes our attempt to construct such a baseline. We reviewed evaluation criteria used to assess cost models and an evaluation framework that was intended to assess the quality of requirements models. We developed an extended evaluation framework and an associated evaluation process that will be used to evaluate our bidding model. Furthermore, we suggest the evaluation framework might be suitable for evaluating other models derived from expert-opinion based influence diagrams. Barbara A. Kitchenham, Lesley Pickard, Stephen G. Linkman, Peter Jones 0002 |
Inf. Softw. Technol. | 1 |
| 2005 | An investigation of software engineering curriculaabstractWe adapted a survey instrument developed by Timothy Lethbridge to assess the extent to which the education delivered by four UK universities matches the requirements of the software industry. We propose a survey methodology that we believe addresses the research question more appropriately than the one used by Lethbridge. In particular, we suggest that restricting the scope of the survey to address the question of whether the curricula for a specific university addressed the needs of its own students, allowed us to identify an appropriate target population. However, our own survey suffered from several problems. In particular the questions used in the survey are not ideal, and the response rate was poor. Although the poor response rate reduces the value of our results, our survey appears to confirm several of Lethbridge's observations with respect to the over-emphasis of mathematical topics and the under-emphasis on business topics. We also have a close agreement with respect to the relative importance of different software engineering topics. However the set of topics, that we found were taught far less than their importance would suggest, were quite different from the topics identified by Lethbridge. Barbara A. Kitchenham, David Budgen, Pearl Brereton, Philip Woodall |
J. Syst. Softw. | 1 |
| 2005 | Erratum to "An empirical study of maintenance and development estimation accuracy" [The Journal of Systems and Software 64 (2002) 57-77]
Barbara A. Kitchenham, Shari Lawrence Pfleeger, Beth McColl, Suzanne Eagan |
J. Syst. Softw. | 1 |
| 2004 | An Exploratory Study of Groupware Support for Distributed Software Architecture Evaluation Process
Muhammad Ali Babar 0001, Barbara A. Kitchenham, Liming Zhu 0001, D. Ross Jeffery |
APSEC | 2 |
| 2004 | Evidence-Based Software EngineeringabstractOur objective is to describe how software engineering might benefit from an evidence-based approach and to identify the potential difficulties associated with the approach. We compared the organisation and technical infrastructure supporting evidence-based medicine (EBM) with the situation in software engineering. We considered the impact that factors peculiar to software engineering (i.e. the skill factor and the lifecycle factor) would have on our ability to practice evidence-based software engineering (EBSE). EBSE promises a number of benefits by encouraging integration of research results with a view to supporting the needs of many different stakeholder groups. However, we do not currently have the infrastructure needed for widespread adoption of EBSE. The skill factor means software engineering experiments are vulnerable to subject and experimenter bias. The lifecycle factor means it is difficult to determine how technologies will behave once deployed. Software engineering would benefit from adopting what it can of the evidence approach provided that it deals with the specific problems that arise from the nature of software engineering. Barbara A. Kitchenham, Tore Dybå, Magne Jørgensen |
ICSE | 1 |
| 2004 | Software Productivity Measurement Using Multiple Size MeasuresabstractProductivity measures based on a simple ratio of product size to project effort assume that size can be determined as a single measure. If there are many possible size measures in a data set and no obvious model for aggregating the measures into a single measure, we propose using the expression AdjustedSize/Effort to measure productivity. AdjustedSize is defined as the most appropriate regression-based effort estimation model, where all the size measures selected for inclusion in the estimation model have a regression parameter significantly different from zero (p<0.05). This productivity measurement method ensures that each project has an expected productivity value of one. Values between zero and one indicate lower than expected productivity, values greater than one indicate higher than expected productivity. We discuss the assumptions underlying this productivity measurement method and present an example of its use for Web application projects. We also explain the relationship between effort prediction models and productivity models. Barbara A. Kitchenham, Emilia Mendes |
IEEE Trans. Software Eng. | 1 |
| 2003 | A Further Empirical Investigation of the Relationship Between MRE and Project Size
Erik Stensrud, Tron Foss, Barbara A. Kitchenham, Ingunn Myrtveit |
Empir. Softw. Eng. | 3 |
| 2003 | A Simulation Study of the Model Evaluation Criterion MMREabstractThe mean magnitude of relative error, MMRE, is probably the most widely used evaluation criterion for assessing the performance of competing software prediction models. One purpose of MMRE is to assist us to select the best model. In this paper, we have performed a simulation study demonstrating that MMRE does not always select the best model. Our findings cast some doubt on the conclusions of any study of competing software prediction models that use MMRE as a basis of model comparison. We therefore recommend not using MMRE to evaluate and compare prediction models. At present, we do not have any universal replacement for MMRE. Meanwhile, we therefore recommend using a combination of theoretical justification of the models that are proposed together with other metrics proposed in this paper. Tron Foss, Erik Stensrud, Barbara A. Kitchenham, Ingunn Myrtveit |
IEEE Trans. Software Eng. | 3 |
| 2003 | Modeling Software Bidding RisksabstractWe discuss a method of developing a software bidding model that allows users to visualize the uncertainty involved in pricing decisions and make appropriate bid/no bid decisions. We present a generic bidding model developed using the modeling method. The model elements were identified after a review of bidding research in software and other industries. We describe the method we developed to validate our model and report the main results of our model validation, including the results of applying the model to four bidding scenarios. Barbara A. Kitchenham, Lesley Pickard, Stephen G. Linkman, Peter Jones 0002 |
IEEE Trans. Software Eng. | 1 |
| 2002 | The question of scale economies in software - why cannot researchers agree?abstractThis paper investigates the different research results obtained when different researchers have investigated the issue of economies and diseconomies of scale in software projects. Although researchers have used broadly similar sets of software project data sets, the results of their analyses and the conclusions they have drawn have differed. The paper highlights methodological differences that have lead to the conflicting results and shows how in many cases the differing results can be reconciled. It discusses the application of econometric concepts such as production frontiers and data envelopment analysis (DEA) to software data sets. It concludes that the assumptions underlying DEA may make it unsuitable for most software datasets but stochastic production frontiers may be relevant. It also raises some statistical issues that suggest testing hypothesis about economies and diseconomies of scale may be much more difficult than has been appreciated. The paper concludes with a plea for agreed standards for research synthesis activities. Barbara A. Kitchenham |
Inf. Softw. Technol. | 1 |
| 2002 | An empirical study of maintenance and development estimation accuracyabstractWe analyzed data from 145 maintenance and development projects managed by a single outsourcing company, including effort and duration estimates, effort and duration actuals, and function points counts. The estimates were made as part of the company’s standard project estimating process that involved producing two or more estimates for each project and selecting one estimate to be the basis of client-agreed budgets. We found that effort estimates chosen as a basis for project budgets were, in general, reasonably good, with 63% of the estimates being within 25% of the actual value, and an average absolute error of 0.26. These estimates were significantly better than regression estimates based on adjusted function points, although the function point models were based on a homogeneous subset of the full data set, and we allowed for the fact that the model parameters changed over time. Furthermore, there was little evidence that the accuracy of the selected estimates was due to their becoming the target values for the project managers. Barbara A. Kitchenham, Shari Lawrence Pfleeger, Beth McColl, Suzanne Eagan |
J. Syst. Softw. | 1 |
| 2002 | Preliminary Guidelines for Empirical Research in Software EngineeringabstractEmpirical software engineering research needs research guidelines to improve the research and reporting processes. We propose a preliminary set of research guidelines aimed at stimulating discussion among software researchers. They are based on a review of research guidelines developed for medical researchers and on our own experience in doing and reviewing software engineering research. The guidelines are intended to assist researchers, reviewers, and meta-analysts in designing, conducting, and evaluating empirical studies. Editorial boards of software engineering journals may wish to use our recommendations as a basis for developing guidelines for reviewers and for framing policies for dealing with the design, data collection, and analysis and reporting of empirical studies. Barbara A. Kitchenham, Shari Lawrence Pfleeger, Lesley Pickard, Peter Jones 0002, David C. Hoaglin, Khaled El Emam, Jarrett Rosenberg |
IEEE Trans. Software Eng. | 1 |
| 2001 | An investigation of coupling, reuse and maintenance in a commercial C++ applicationabstractThis paper describes an investigation into the use of coupling complexity metrics to obtain early indications of various properties of a system of C++ classes. The properties of interest are: (i) the potential reusability of a class and (ii) the likelihood that a class will be affected by maintenance changes made to the overall system. The study indicates that coupling metrics can provide useful indications of both reusable classes and of classes that may have a significant influence on the effort expended during system maintenance and testing. F. George Wilkie, Barbara A. Kitchenham |
Inf. Softw. Technol. | 2 |
| 2001 | Modeling Software Measurement DataabstractThis paper proposes a method for specifying models of software data sets in order to capture the definitions and relationships among software measures. We believe a method of defining software data sets is necessary to ensure that software data are trustworthy. Software companies introducing a measurement program need to establish procedures to collect and store trustworthy measurement data. Without appropriate definitions it is difficult to ensure data values are repeatable and comparable. Software metrics researchers need to maintain collections of software data sets. Such collections allow researchers to assess the generality of software engineering phenomena. Without appropriate safeguards, it is difficult to ensure that data from different sources are analyzed correctly. These issues imply the need for a standard method of specifying software data sets so they are fully documented and can be exchanged with confidence. We suggest our method of defining data sets can be used as such a standard. We present our proposed method in terms of a conceptual entity-relationship data model that allows complex software data sets to be modeled and their data values stored. The standard can, therefore, contribute both to the definition of a company measurement program and to the exchange of data sets among researchers. Barbara A. Kitchenham, Robert T. Hughes, Stephen G. Linkman |
IEEE Trans. Software Eng. | 1 |
| 2000 | An evaluation of the business object approach to software developmentabstractIn this paper, we report the result of an evaluation of the use of business objects and business components for developing business application software. This evaluation was a replicated product case study in which a part of an existing product was re-implemented using an UML-based development process. In order to assess the impact of re-use, a second related product was implemented using the new technology. We found that producing software from scratch using UML was less productive during the development lifecycle, but productivity improved when substantial reuse (48%) was achieved. Time to market was not much affected by the new technology but was greatly improved when substantial reuse was achieved. Defect rates appeared substantially lower for the new technology irrespective of reuse levels. The technology also had other benefits including provision of documentation and less reliance on individual members of staff. Manolis Tsagias, Barbara A. Kitchenham |
J. Syst. Softw. | 2 |
| 2000 | Coupling measures and change ripples in C++ application softwareabstractThis paper describes an investigation into the effects of class couplings on changes made to a commercial C++ application over a period of 212 yr. The Chidamber and Kemerer CBO metric is used to measure class couplings within the application and its limitations are identified. Through an in-depth study of the ripple effects of changes to the source code, practical insight into the nature and extent of software couplings is provided. F. George Wilkie, Barbara A. Kitchenham |
J. Syst. Softw. | 2 |
| 1999 | Towards an ontology of software maintenanceabstractWe suggest that empirical studies of maintenance are difficult to understand unless the context of the study is fully defined. We developed a preliminary ontology to identify a number of factors that influence maintenance. The purpose of the ontology is to identify factors that would affect the results of empirical studies. We present the ontology in the form of a UML model. Using the maintenance factors included in the ontology, we define two common maintenance scenarios and consider the industrial issues associated with them. Copyright © 1999 John Wiley & Sons, Ltd. Barbara A. Kitchenham, Guilherme Horta Travassos, Anneliese Amschler Andrews, Frank Niessink, Norman F. Schneidewind, Janice Singer, Shingo Takada 0001, Risto Vehvilainen |
J. Softw. Maintenance Res. Pract. | 1 |
| 1999 | Comments on: Evaluating Alternative Software Production FunctionsabstractSoftware development projects are notorious for cost overruns and schedule delays. While dozens of software cost models have been proposed, few of them seem to have any degree of consistent accuracy. One major factor contributing to this persistent and widespread problem is an inadequate understanding of the real behavior of software development processes. We believe that software development could be studied as an economic production process and that established economic theories and methods could be used to develop and validate software production and cost models. We present the results of evaluating four alternative software production models using the P-test, a statistical procedure developed specifically for testing the truth of a hypothesis in the presence of alternatives in econometric studies. We found that the truth of the widely used Cobb-Douglas type of software production and cost models (e.g., COCOMO) cannot be maintained in the presence of quadratic or translog models. Overall, the quadratic software production function is shown to be the most plausible model for representing software production processes. Limitations of this study and future directions are also discussed. Lesley Pickard, Barbara A. Kitchenham, Peter Jones 0002 |
IEEE Trans. Software Eng. | 2 |
| 1998 | Combining empirical results in software engineeringabstractIn this paper we investigate the techniques used in medical research to combine results from independent empirical studies of a particular phenomenon: meta-analysis and vote-counting. We use an example to illustrate the benefits and limitations of each technique and to indicate the criteria that should be used to guide your choice of technique. Meta-analysis is appropriate for homogeneous studies when raw data or quantitative summary information, e.g. correlation coefficient, are available. It can also be used for heterogeneous studies where the cause of the heterogeneity is due to well-understood partitions in the subject population. In other circumstances, meta-analysis is usually invalid. Although intuitively appealing, vote-counting has a number of serious limitations and should usually be avoided. We suggest that combining study results is unlikely to solve all the problems encountered in empirical software engineering studies, but some of the infrastructure and controls used by medical researchers to improve the quality of their empirical studies would be useful in the field of software engineering. Lesley Pickard, Barbara A. Kitchenham, Peter Jones 0002 |
Inf. Softw. Technol. | 2 |
| 1998 | A Procedure for Analyzing Unbalanced DatasetsabstractThis paper describes a procedure for analyzing unbalanced datasets that include many nominal- and ordinal-scale factors. Such datasets are often found in company datasets used for benchmarking and productivity assessment. The two major problems caused by lack of balance are that the impact of factors can be concealed and that spurious impacts can be observed. These effects are examined with the help of two small artificial datasets. The paper proposes a method of forward pass residual analysis to analyze such datasets. The analysis procedure is demonstrated on the artificial datasets and then applied to the COCOMO dataset. The paper ends with a discussion of the advantages and limitations of the analysis procedure. Barbara A. Kitchenham |
IEEE Trans. Software Eng. | 1 |
| 1997 | Evaluation and assessment in software engineering
Barbara A. Kitchenham, Pearl Brereton, David Budgen, Stephen G. Linkman, Vicki L. Almstrum, Shari Lawrence Pfleeger |
Inf. Softw. Technol. | 1 |
| 1997 | Experiences introducing a measurement programabstractMeasurement is an integral part of total quality management and process improvement strategies. This paper describes our experiences using the Goal-Question-Metric (GQM) paradigm to help design a company-wide measurement program for Engineering Ingegneria S.p.A., an Italian software house. The introduction of the measurement program was supported by the Commission of the European Communities within the European Software and Systems Initiative (ESSI) as a Process Improvement Experiment (PIE). We found it necessary to supplement GQM into two ways. Firstly, we defined our measures rigorously in terms of entities, attributes, units and counting rules. Secondly, the original GQM plan was subject to an independent review. The most critical problem identified by the review was that the GQM plan identified too many productivity factors for any statistical analysis to handle concurrently. In order to address this issue, we developed an analysis technique based on a step-wise analysis of residuals. This has allowed us to identify the main factors affecting productivity and effort. Stefano De Panfilis, Barbara A. Kitchenham, N. Morfuni |
Inf. Softw. Technol. | 2 |
| 1997 | The SQUID approach to defining a quality model
Barbara A. Kitchenham, Stephen G. Linkman, Alberto Pasquini, Vincenzo Nanni |
Softw. Qual. J. | 1 |
| 1997 | Reply to: Comments on "Toward a Framework for Software Measurement Validation"abstractare glad that members of the software engineering communityhave taken-up our suggestion that metric validation needs publicdebate. With respect to specific issues raised by Basili et al. in theirletter, clearly, we disagree with their view that we are seriouslyimpeding the appropriate definition of attributes; we are glad tohave the opportunity to clarify some of our arguments. Barbara A. Kitchenham, Shari Lawrence Pfleeger, Norman E. Fenton |
IEEE Trans. Software Eng. | 1 |
| 1996 | Effort Estimation Using Analogy
Martin J. Shepperd, Chris Schofield, Barbara A. Kitchenham |
ICSE | 3 |
| 1995 | Towards a Framework for Software Measurement ValidationabstractIn this paper we propose a framework for validating software measurement. We start by defining a measurement structure model that identifies the elementary component of measures and the measurement process, and then consider five other models involved in measurement: unit definition models, instrumentation models, attribute relationship models, measurement protocols and entity population models. We consider a number of measures from the viewpoint of our measurement validation framework and identify a number of shortcomings; in particular we identify a number of problems with the construction of function points. We also compare our view of measurement validation with ideas presented by other researchers and identify a number of areas of disagreement. Finally, we suggest several rules that practitioners and researchers can use to avoid measurement problems, including the use of measurement vectors rather than artificially contrived scalars. Barbara A. Kitchenham, Shari Lawrence Pfleeger, Norman E. Fenton |
IEEE Trans. Software Eng. | 1 |
| 1993 | Inter-item Correlations among Function Points
Barbara A. Kitchenham, Kari Känsälä |
ICSE | 1 |
| 1993 | Letter to the editor
Barbara A. Kitchenham |
Inf. Softw. Technol. | 1 |
| 1992 | Empirical studies of assumptions that underlie software cost-estimation models
Barbara A. Kitchenham |
Inf. Softw. Technol. | 1 |
| 1991 | Validating Software Measures
Norman E. Fenton, Barbara A. Kitchenham |
Softw. Test. Verification Reliab. | 2 |
| 1988 | An evaluation of software structure metricsabstractEvaluates some software design metrics, based on the information flow metrics of S. Henry and D. Kafura (1981, 1984), using data from a communications system. The ability of the design metrics to identify change-prone, error-prone, and complex programs was contrasted with that of simple code metrics. It was found that the design metrics were not as good at identifying change-prone, fault-prone, and complex programs as simple code metrics (i.e. lines of code and number of branches). It was also observed that the compound metrics, built up from several different basic counts, can obscure underlying effects, and thus make it more difficult to use metrics constructively.> Barbara A. Kitchenham |
COMPSAC | 1 |
| 1985 | A Comparison of Cost Estimation Tools (Panel)
Barbara A. Kitchenham, Howard A. Rubin |
ICSE | 1 |
| 1985 | Software project development cost estimationabstractThis paper reports the results of an empirical investigation of the relationships between effort expended, time scales, and project size for software project development. The observed relationships were compared with those predicted by Lawrence Putnam's Rayleigh curve model and Barry Boehm's COCOMO model. The results suggested that although the form of the basic empirical relationships were consistent with the cost models, the COCOMO model was a poor estimator of cost for the current data set and the data did not follow the Rayleigh curve suggested by Putnam. However, the results did suggest that it was possible to develop cost models tailored to a particular environment and to improve the precision of the models as they are used during the development cycle by including additional information such as the known effort for the early development phases. The paper finishes by discussing some of the problems involved in developing useful cost models. Barbara A. Kitchenham, N. R. Taylor |
J. Syst. Softw. | 1 |