VLDB 2026 Research / reviewers in the wild / expert
Lech Madeyski
dblp:74/4587
· DBLP profile ↗
40ranked-venue papers
17as first author
20since 2021 · last 2026
0000-0003-3907-3357ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 33 · 17 first-author · 16 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Test case prioritization: A systematic review using snowballing and TCPFramework with approach combinatorsabstractContext: Test case prioritization (TCP) is a technique widely used by software development organizations to accelerate regression testing. Objective: We aim to systematize existing TCP knowledge and to propose and empirically evaluate a new TCP approach. Methods: We conduct a systematic review (SR) using snowballing on TCP, implement a comprehensive platform for TCP research (TCPFramework), analyze existing evaluation metrics and propose two new ones ( r APFD C and ATR), and develop a family of ensemble TCP methods called approach combinators. Results: The SR helped identify 324 studies related to TCP. The techniques proposed in our study were evaluated on the RTPTorrent dataset, consistently outperforming their base approaches across the majority of subject programs, and achieving performance comparable to the current state of the art for heuristical algorithms (in terms of r APFD C , NTR, and ATR), while using a distinct approach. Conclusion: The proposed methods can be used efficiently for TCP, reducing the time spent on regression testing by up to 2.7%. Approach combinators offer significant potential for improvements in future TCP research, due to their composability. Tomasz Chojnacki, Lech Madeyski |
Inf. Softw. Technol. | 2 |
| 2026 | How vulnerability explanations help software practitioners confirm and fix code vulnerabilitiesabstractContext: Most current code vulnerability detection tools provide only a binary classification (vulnerable/non-vulnerable) with little to no additional context. This paper explores the impact of providing explanations for vulnerabilities alongside code labelled as vulnerable. Objective: We investigate the influence of explanations on the ability of software practitioners to confirm such labelled code as actually vulnerable (i.e., a true positive vulnerability) and to fix such vulnerable code correctly. Method: We surveyed 99 software practitioners to establish their use of code-vulnerability detection tools and to evaluate the impact of explanations on their behaviour towards code labelled as vulnerable in a series of coding exercises. Participants were presented with four forms of explanation: vulnerable lines , vulnerability type , short-form text , and long-form text . Results: Software practitioners performed better at confirming and fixing code vulnerabilities when presented with any of the four forms of explanation. Although practitioners stated a preference for long-form text explanations, they achieved the highest confirmation and fixing performance with short-form text explanations. Practitioners also indicated willingness to accept modest drops in detection precision and recall if richer explanations were provided, and their preferences for explanation types and performance trade-offs varied according to where a detection tool is used in the software-development pipeline. Conclusions: Vulnerability-detection and prediction tools should provide explanatory output and allow different explanation types tailored to their deployment stage in the development workflow. Few current tools provide any explanations, and none identified in this study provide text-based explanations. Fahad Al Debeyan, Tracy Hall, Lech Madeyski, Emily Winter 0001 |
Inf. Softw. Technol. | 3 |
| 2026 | LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviewsabstractContext: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matrix metrics used to evaluate them can mislead under the imbalanced, cost-asymmetric conditions of screening. Objective: We develop and justify LLM4SCREENLIT—practical recommendations for researchers conducting LLM-screening evaluations and for editors and reviewers assessing such studies—differentiated by study type (retrospective benchmarking vs. deployment for a specific SR). Method: Using Delgado-Chaves et al. (2025), an 18-LLM benchmark across three biomedical SRs, as a motivating example, we reviewed 28 additional papers and extracted their reported metrics. We propose a Weighted Matthews Correlation Coefficient (WMCC) that integrates MCC’s chance-correction with asymmetric misclassification costs, and validated it on three software-engineering (SE) reanalyses (Felizardo et al. 2024; Syriani et al. 2024; Huotala et al. 2025), the largest covering 9 LLMs × 24 SE secondary studies (34,528 articles). Results: Across the 29 papers, only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost. In the largest SE reanalysis, MCC and WMCC disagree on the best LLM in 55% of evaluable studies; in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%. Sensitivity analysis (median crossover at w ≈ 2 . 7 , all < 7 ) supports w = 10 as a conservative default. Conclusions: SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available. Editors and reviewers should require these elements as routine. Extension to full-text screening and data extraction is principled but pending empirical validation. Lech Madeyski, Barbara A. Kitchenham, Martin J. Shepperd |
Inf. Softw. Technol. | 1 |
| 2025 | Predicting test failures induced by software defects: A lightweight alternative to software defect prediction and its industrial applicationabstractContext: Machine Learning Software Defect Prediction (ML SDP) is a promising method to improve the quality and minimise the cost of software development. Objective: We aim to: (1) apropose and develop a Lightweight Alternative to SDP (LA2SDP) that predicts test failures induced by software defects to allow pinpointing defective software modules thanks to available mapping of predicted test failures to past defects and corrected modules, (2) preliminary evaluate the proposed method in a real-world Nokia 5G scenario. Method: We train machine learning models using test failures that come from confirmed software defects already available in the Nokia 5G environment. We implement LA2SDP using five supervised ML algorithms, together with their tuned versions, and use eXplainable AI (XAI) to provide feedback to stakeholders and initiate quality improvement actions. Results: We have shown that LA2SDP is feasible in vivo using test failure-to-defect report mapping readily available within the Nokia 5G system-level test process, achieving good predictive performance . Specifically, CatBoost Gradient Boosting turned out to perform the best and achieved satisfactory Matthew’s Correlation Coefficient (MCC) results for our feasibility study . Conclusions: Our efforts have successfully defined, developed, and validated LA2SDP, using the sliding and expanding window approaches on an industrial data set. Lech Madeyski, Szymon Stradowski |
J. Syst. Softw. | 1 |
| 2025 | "Your AI is impressive, but my code does not have any bugs" managing false positives in industrial contextsabstractContext “Your AI is impressive, but my code does not contain any bugs”— such a statement from a software developer is the antithesis of a quality mindset and open communication. What makes it worse is that it is oftentimes true. Objective This paper analyses false positives' impact and related challenges in machine learning software defect prediction and describes the mitigation possibilities. Methods We propose a broad-picture perspective on dealing with false positive predictions based on what we learned from our industrial implementation study in Nokia 5G. Results Accordingly, we draw a new direction in transitioning defect prediction into a well-established industry practice, as well as highlight potential emerging topics in predictive software engineering. Conclusion Increasing human buy-in and the business impact of predictions significantly improves the chances of future software defect prediction industry adoptions to succeed. Szymon Stradowski, Lech Madeyski |
Sci. Comput. Program. | 2 |
| 2024 | An Empirical Analysis of the Usage of Requirements Attributes in Requirements Engineering Research and Practice
Krzysztof Wnuk, Lech Madeyski, Waleed Abdeen, Sneha Penmetsa, Navya Lingampalli |
ICCCI (2) | 2 |
| 2024 | Recommendations for analysing and meta-analysing small sample size software engineering experimentsabstractAbstract Context Software engineering (SE) experiments often have small sample sizes. This can result in data sets with non-normal characteristics, which poses problems as standard parametric meta-analysis, using the standardized mean difference (StdMD) effect size, assumes normally distributed sample data. Small sample sizes and non-normal data set characteristics can also lead to unreliable estimates of parametric effect sizes. Meta-analysis is even more complicated if experiments use complex experimental designs, such as two-group and four-group cross-over designs, which are popular in SE experiments. Objective Our objective was to develop a validated and robust meta-analysis method that can help to address the problems of small sample sizes and complex experimental designs without relying upon data samples being normally distributed. Method To illustrate the challenges, we used real SE data sets. We built upon previous research and developed a robust meta-analysis method able to deal with challenges typical for SE experiments. We validated our method via simulations comparing StdMD with two robust alternatives: the probability of superiority ( $$\hat{p}$$ p ^ ) and Cliffs’ d. Results We confirmed that many SE data sets are small and that small experiments run the risk of exhibiting non-normal properties, which can cause problems for analysing families of experiments. For simulations of individual experiments and meta-analyses of families of experiments, $$\hat{p}$$ p ^ and Cliff’s d consistently outperformed StdMD in terms of negligible small sample bias. They also had better power for log-normal and Laplace samples, although lower power for normal and gamma samples. Tests based on $$\hat{p}$$ p ^ always had better or equal power than tests based on Cliff’s d, and across all but one simulation condition, $$\hat{p}$$ p ^ Type 1 error rates were less biased. Conclusions Using $$\hat{p}$$ p ^ is a low-risk option for analysing and meta-analysing data from small sample-size SE randomized experiments. Parametric methods are only preferable if you have prior knowledge of the data distribution. Barbara A. Kitchenham, Lech Madeyski |
Empir. Softw. Eng. | 2 |
| 2024 | The impact of hard and easy negative training data on vulnerability prediction performanceabstractVulnerability prediction models have been shown to perform poorly in the real world. We examine how the composition of negative training data influences vulnerability prediction model performance. Inspired by other disciplines (e.g. image processing), we focus on whether distinguishing between negative training data that is ‘easy’ to recognise from positive data (very different from positive data) and negative training data that is ‘hard’ to recognise from positive data (very similar to positive data) impacts on vulnerability prediction performance. We use a range of popular machine learning algorithms, including deep learning, to build models based on vulnerability patch data curated by Reis and Abreu, as well as the MSR dataset. Our results suggest that models trained on higher ratios of easy negatives perform better, plateauing at 15 easy negatives per positive instance. We also report that different ML algorithms work better based on the negative sample used. Overall, we found that the negative sampling approach used significantly impacts model performance, potentially leading to overly optimistic results. The ratio of ‘easy’ versus ‘hard’ negative training data should be explicitly considered when building vulnerability prediction models for the real world. Fahad Al Debeyan, Lech Madeyski, Tracy Hall, David Bowes |
J. Syst. Softw. | 2 |
| 2023 | Bridging the Gap Between Academia and Industry in Machine Learning Software Defect Prediction: Thirteen ConsiderationsabstractThis experience paper describes thirteen considerations for implementing machine learning software defect prediction (ML SDP) in vivo. Specifically, we provide the following report on the ground of the most important observations and lessons learned gathered during a large-scale research effort and introduction of ML SDP to the system-level testing quality assurance process of one of the leading telecommunication vendors in the world — Nokia. We adhere to a holistic and logical progression based on the principles of the business analysis body of knowledge: from identifying the need and setting requirements, through designing and implementing the solution, to profitability analysis, stakeholder management, and handover. Conversely, for many years, industry adoption has not kept up the pace of academic achievements in the field, despite promising potential to improve quality and decrease the cost of software products for many companies worldwide. Therefore, discussed considerations hopefully help researchers and practitioners bridge the gaps between academia and industry. Szymon Stradowski, Lech Madeyski |
ASE | 2 |
| 2023 | Continuous build outcome prediction: an experimental evaluation and acceptance modelling
Marcin Kawalerowicz, Lech Madeyski |
Appl. Intell. | 2 |
| 2023 | Detecting code smells using industry-relevant data
Lech Madeyski, Tomasz Lewowski |
Inf. Softw. Technol. | 1 |
| 2023 | Exploring the challenges in software testing of the 5G system at Nokia: A surveyabstractThe ever-growing size and complexity of industrial software products pose significant quality assurance challenges to engineering researchers and practitioners, despite the constant effort to increase knowledge and improve the processes. 5G technology developed by Nokia is one example of such a grand and highly complex system with improvement potential. The following paper provides an overview of the current quality assurance processes used by Nokia to develop the 5G technology and provides insight into the most prominent challenges by an evaluation of perceived importance, urgency, and difficulty to understand the future opportunities. Nokia mode of operation, briefly introduced in this paper, has been subjected to extensive analysis by a selected group of experienced test-oriented professionals to define the most critical areas of concern. Secondly, the identified problems were evaluated by Nokia gNB system-level test professionals in a dedicated survey. The questionnaire was completed by 312 out of 2935 (10.63%) possible respondents. The challenges are seen as the most important and urgent: customer scenario testing, performance testing, and competence ramp-up. Challenges seen as the most difficult to solve are low occurrence failures, hidden feature dependencies, and hardware configuration-specific problems. Our research identified several improvement areas in the quality assurance processes used to develop the 5G technology by determining the most important and urgent problems that at the same time have a low perceived difficulty. Such initiatives are attractive from a business perspective. On the other hand, challenges seen as the most impactful yet difficult may be of interest to the academic research community. Szymon Stradowski, Lech Madeyski |
Inf. Softw. Technol. | 2 |
| 2023 | Machine learning in software defect prediction: A business-driven systematic mapping study
Szymon Stradowski, Lech Madeyski |
Inf. Softw. Technol. | 2 |
| 2023 | Industrial applications of software defect prediction using machine learning: A business-driven systematic literature reviewabstractMachine learning software defect prediction is a promising field of software engineering, attracting a great deal of attention from the research community; however, its industry application tents to lag behind academic achievements. This study is part of a larger project focused on improving the quality and minimising the cost of software testing of the 5G system at Nokia, and aims to evaluate the business applicability of machine learning software defect prediction and gather lessons learnt. The systematic literature review was conducted on journal and conference papers published between 2015 and 2022 in popular online databases (ACM, IEEE, Springer, Scopus, Science Direct, and Google Scholar). A quasi-gold standard procedure was used to validate the search, and SEGRESS guidelines were used for transparency, reporting, and replicability. We have selected and analysed 32 publications out of 397 found by our automatic search (and seven by snowballing). We have identified highly relevant evidence of methods, features, frameworks, and datasets used. However, we found a minimal emphasis on practical lessons learnt and cost consciousness — both vital from a business perspective. Even though the number of machine learning software defect prediction studies validated in the industry is increasing (and we were able to identify several excellent papers on studies performed in vivo), there is still not enough practical focus on the business aspects of the effort that would help bridge the gap between the needs of the industry and academic research. Szymon Stradowski, Lech Madeyski |
Inf. Softw. Technol. | 2 |
| 2023 | How Should Software Engineering Secondary Studies Include Grey Material?abstractContext: Recent papers have proposed the use ofgrey literature(GL) and multivocal reviews. These papers have raised issues about the practices used for systematic reviews (SRs) in software engineering (SE) and suggested that there should be changes to the current SR guidelines.Objective: To investigate whether current SR guidelines need to be changed to support GL and multivocal reviews.Method: We discuss the definitions of GL and the importance of GL and of industry-based field studies in SE SRs. We identify properties of SRs that constrain the material used in SRs: a) the nature of primary studies; b) the requirements of SRs to be auditable, traceable, and reproducible; and explain why these requirements restrict the use of blogs in SRs.Results: SR guidelines have always considered GL as a possible source of primary studies and have never supported exclusion of field studies that incorporate the practitioners’ viewpoint. However, the concept of GL, which was meant to refer to documents that were not formally published, is now being extended to information from sources such as blogs/tweets/Q&A posts. Thus, it might seem that SRs do not make full use of GL because they do not include such information. However, the unit of analysis for an SR is the primary study. Thus, it is not thesourcebut thetypeof information that is important. Any report describing a rigorous empirical evaluation is a candidate primary study. Whether it is actually included in an SR depends on the SR eligibility criteria. However, any study that cannot be guaranteed to be publicly available in the long term should not be used as a primary study in an SR. This does not prevent such information from being aggregated in surveys of social media and used in the context of evidence-based software engineering (EBSE).Conclusions: Current guidelines for SRs do not require extensions, but their scope needs to be better defined. SE researchers require guidelines for analysing social media posts (e.g., blogs, tweets, vlogs), but these should be based on qualitative primary (not secondary) study guidelines. SE researchers can use mixed-methods SRs and/or the fourth step of EBSE to incorporate findings from social media surveys with those from SRs and to develop industry-relevant recommendations. Barbara A. Kitchenham, Lech Madeyski, David Budgen |
IEEE Trans. Software Eng. | 2 |
| 2023 | SEGRESS: Software Engineering Guidelines for REporting Secondary StudiesabstractContext: Several tertiary studies have criticized the reporting of software engineering secondary studies.Objective: Our objective is to identify guidelines for reporting software engineering (SE) secondary studies which would address problems observed in the reporting of software engineering systematic reviews (SRs).Method: We review the criticisms of SE secondary studies and identify the major areas of concern. We assess the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) statement as a possible solution to the need for SR reporting guidelines, based on its status as the reporting guideline recommended by the Cochrane Collaboration whose SR guidelines were a major input to the guidelines developed for SE. We report its advantages and limitations in the context of SE secondary studies. We also assess reporting guidelines for mapping studies and qualitative reviews, and compare their structure and content with that of PRISMA 2020.Results: Previous tertiary studies confirm that reports of secondary studies are of variable quality. However,ad hocrecommendations that amend reporting standards may result in unnecessary duplication of text. We confirm that the PRISMA 2020 statement addresses SE reporting problems, but is mainly oriented to quantitative reviews, mixed-methods reviews and meta-analyses. However, we show that the PRISMA 2020 item definitions can be extended to cover the information needed to report mapping studies and qualitative reviews.Conclusions: In this paper and its Supplementary Material, we present and illustrate an integrated set of guidelines called SEGRESS (Software Engineering Guidelines for REporting Secondary Studies), suitable for quantitative systematic reviews (building upon PRISMA 2020), mapping studies (PRISMA-ScR), and qualitative reviews (ENTREQ and RAMESES), that addresses reporting problems found in current SE SRs. Barbara A. Kitchenham, Lech Madeyski, David Budgen |
IEEE Trans. Software Eng. | 2 |
| 2022 | How far are we from reproducible research on code smell detection? A systematic literature reviewabstractCode smells are symptoms of wrong design decisions or coding shortcuts that may increase defect rate and decrease maintainability. Research on code smells is accelerating, focusing on code smell detection and using code smells as defect predictors. Recent research shows that even between software developers, agreement on what constitutes a code smell is low, but several publications claim the high performance of detection algorithms—which seems counterintuitive, considering that algorithms should be taught on data labeled by developers. This paper aims to investigate the possible reasons for the inconsistencies between studies in the performance of applied machine learning algorithms compared to developers. It focuses on the reproducibility of existing studies. A systematic literature review was performed among conference and journal articles published between 1999 and 2020 to assess the state of reproducibility of the research performed in those papers. A quasi-gold standard procedure was used to validate the search. Modeling process descriptions, reproduction scripts, data sets, and techniques used for their creation were analyzed. We obtained data from 46 publications. 22 of them contained a detailed description of the modeling process, 17 included any reproduction data (data set, results, or scripts) and 15 used existing data sets. In most of the publications, analyzed projects were hand-picked by the researchers. Most studies do not include any form of an online reproduction package, although this has started to change recently—8% of analyzed studies published before 2018 included a full reproduction package, compared to 22% in years 2018–2019. Ones that do include a package usually use a research group website or even a personal one. Dedicated archives are still rarely used for data packages. We recommend that researchers include complete reproduction packages for their studies and use well-established research data archives instead of their own websites. Tomasz Lewowski, Lech Madeyski |
Inf. Softw. Technol. | 2 |
| 2022 | The Importance of the Correlation in Crossover ExperimentsabstractContext:In empirical software engineering, crossover designs are popular for experiments comparing software engineering techniques that must be undertaken by human participants. However, their value depends on the correlation ($r$) between the outcome measures on the same participants. Software engineering theory emphasizes the importance of individual skill differences, so we would expect the values of$r$to be relatively high. However, few researchers have reported the values of$r$.Goal:To investigate the values of$r$found in software engineering experiments.Method:We undertook simulation studies to investigate the theoretical and empirical properties of$r$. Then we investigated the values of$r$observed in 35 software engineering crossover experiments.Results:The level of$r$obtained by analysing our 35 crossover experiments was small. Estimates based on means, medians, and random effect analysis disagreed but were all between 0.2 and 0.3. As expected, our analyses found large variability among the individual$r$estimates for small sample sizes, but no indication that$r$estimates were larger for the experiments with larger sample sizes that exhibited smaller variability.Conclusions:Low observed$r$values cast doubts on the validity of crossover designs for software engineering experiments. However, if the cause of low$r$values relates to training limitations or toy tasks, this affectsallSoftware Engineering (SE) experiments involving human participants. For all human-intensive SE experiments, we recommend more intensive training and then tracking the improvement of participants as they practice using specific techniques, before formally testing the effectiveness of the techniques. Barbara A. Kitchenham, Lech Madeyski, Giuseppe Scanniello, Carmine Gravino |
IEEE Trans. Software Eng. | 2 |
| 2021 | Continuous Build Outcome Prediction: A Small-N Experiment in Settings of a Real Software Project
Marcin Kawalerowicz, Lech Madeyski |
IEA/AIE (2) | 2 |
| 2021 | Jaskier: A Supporting Software Tool for Continuous Build Outcome Prediction Practice
Marcin Kawalerowicz, Lech Madeyski |
IEA/AIE (2) | 2 |
| 2020 | MLCQ: Industry-Relevant Code Smell Data SetabstractContext Research on code smells accelerates and there are many studies that discuss them in the machine learning context. However, while data sets used by researchers vary in quality, all which we encountered share visible shortcomings---data sets are gathered from a rather small number of often outdated projects by single individuals whose professional experience is unknown. Lech Madeyski, Tomasz Lewowski |
EASE | 1 |
| 2020 | Meta-analysis for families of experiments in software engineering: a systematic review and reproducibility and validity assessmentabstractPrevious studies have raised concerns about the analysis and meta-analysis of crossover experiments and we were aware of several families of experiments that used crossover designs and meta-analysis. To identify families of experiments that used meta-analysis, to investigate their methods for effect size construction and aggregation, and to assess the reproducibility and validity of their results. We performed a systematic review (SR) of papers reporting families of experiments in high quality software engineering journals, that attempted to apply meta-analysis. We attempted to reproduce the reported meta-analysis results using the descriptive statistics and also investigated the validity of the meta-analysis process. Out of 13 identified primary studies, we reproduced only five. Seven studies could not be reproduced. One study which was correctly analyzed could not be reproduced due to rounding errors. When we were unable to reproduce results, we provide revised meta-analysis results. To support reproducibility of analyses presented in our paper, it is complemented by the reproducer R package. Meta-analysis is not well understood by software engineering researchers. To support novice researchers, we present recommendations for reporting and meta-analyzing families of experiments and a detailed example of how to analyze a family of 4-group crossover experiments. Barbara A. Kitchenham, Lech Madeyski, Pearl Brereton |
Empir. Softw. Eng. | 2 |
| 2019 | Problems with Statistical Practice in Human-Centric Software Engineering ExperimentsabstractBackground Examples of questionable statistical practice, when published in high quality software engineering (SE) journals, may lead to novice researchers adopting incorrect statistical practices. Barbara A. Kitchenham, Lech Madeyski, Pearl Brereton |
EASE | 2 |
| 2018 | Effect sizes and their variance for AB/BA crossover design studiesabstractWe addressed the issues related to repeated measures experimental design such as an AB/BA crossover design that have been neither discussed nor addressed in the software engineering literature. Lech Madeyski, Barbara A. Kitchenham |
ICSE | 1 |
| 2018 | Effect sizes and their variance for AB/BA crossover design studiesabstractVegas et al. IEEE Trans Softw Eng 42(2):120:135 (2016) raised concerns about the use of AB/BA crossover designs in empirical software engineering studies. This paper addresses issues related to calculating standardized effect sizes and their variances that were not addressed by the Vegas et al.’s paper. In a repeated measures design such as an AB/BA crossover design each participant uses each method. There are two major implication of this that have not been discussed in the software engineering literature. Firstly, there are potentially two different standardized mean difference effect sizes that can be calculated, depending on whether the mean difference is standardized by the pooled within groups variance or the within-participants variance. Secondly, as for any estimated parameters and also for the purposes of undertaking meta-analysis, it is necessary to calculate the variance of the standardized mean difference effect sizes (which is not the same as the variance of the study). We present the model underlying the AB/BA crossover design and provide two examples to demonstrate how to construct the two standardized mean difference effect sizes and their variances, both from standard descriptive statistics and from the outputs of statistical software. Finally, we discuss the implication of these issues for reporting and planning software engineering experiments. In particular we consider how researchers should choose between a crossover design or a between groups design. Lech Madeyski, Barbara A. Kitchenham |
Empir. Softw. Eng. | 1 |
| 2018 | Introduction to the special section on Enhancing Credibility of Empirical Software Engineering
Lech Madeyski, Barbara A. Kitchenham, Krzysztof Wnuk |
Inf. Softw. Technol. | 1 |
| 2017 | Continuous defect prediction: the idea and a related datasetabstractWe would like to present the idea of our Continuous Defect Prediction (CDP) research and a related dataset that we created and share. Our dataset is currently a set of more than 11 million data rows, representing files involved in Continuous Integration (CI) builds, that synthesize the results of CI builds with data we mine from software repositories. Our dataset embraces 1265 software projects, 30,022 distinct commit authors and several software process metrics that in earlier research appeared to be useful in software defect prediction. In this particular dataset we use TravisTorrent as the source of CI data. TravisTorrent synthesizes commit level information from the Travis CI server and GitHub open-source projects repositories. We extend this data to a file change level and calculate the software process metrics that may be used, for example, as features to predict risky software changes that could break the build if committed to a repository with CI enabled. Lech Madeyski, Marcin Kawalerowicz |
MSR | 1 |
| 2017 | Robust Statistical Methods for Empirical Software EngineeringabstractThere have been many changes in statistical theory in the past 30 years, including increased evidence that non-robust methods may fail to detect important results. The statistical advice available to software engineering researchers needs to be updated to address these issues. This paper aims both to explain the new results in the area of robust analysis methods and to provide a large-scale worked example of the new methods. We summarise the results of analyses of the Type 1 error efficiency and power of standard parametric and non-parametric statistical tests when applied to non-normal data sets. We identify parametric and non-parametric methods that are robust to non-normality. We present an analysis of a large-scale software engineering experiment to illustrate their use. We illustrate the use of kernel density plots, and parametric and non-parametric methods using four different software engineering data sets. We explain why the methods are necessary and the rationale for selecting a specific analysis. We suggest using kernel density plots rather than box plots to visualise data distributions. For parametric analysis, we recommend trimmed means, which can support reliable tests of the differences between the central location of two or more samples. When the distribution of the data differs among groups, or we have ordinal scale data, we recommend non-parametric methods such as Cliff’s δ or a robust rank-based ANOVA-like method. Barbara A. Kitchenham, Lech Madeyski, David Budgen, Jacky W. Keung, Pearl Brereton, Stuart M. Charters, Shirley Gibbs, Amnart Pohthong |
Empir. Softw. Eng. | 2 |
| 2016 | Higher Order Mutation Testing to Drive Development of New Test Cases: An Empirical Comparison of Three Strategies
Quang-Vu Nguyen 0001, Lech Madeyski |
ACIIDS (1) | 2 |
| 2016 | On the Relationship Between the Order of Mutation Testing and the Properties of Generated Higher Order Mutants
Quang-Vu Nguyen 0001, Lech Madeyski |
ACIIDS (1) | 2 |
| 2016 | Empirical Evaluation of Multiobjective Optimization Algorithms Searching for Higher Order MutantsabstractFirst order mutation testing is used to evaluate the quality of a given set of test cases by inserting single changes into the program under test to produce first order mutants (FOMs) of the original program, and then checking whether tests are good enough to detect the artificially injected defects. However, mutation testing is not yet widely used due to the problems of a large number of generated mutants and limited realism of introduced changes that do not necessarily reflect real software defects. Furthermore, many of the generated mutants are equivalent, i.e., they keep the program semantics unchanged and, thus, cannot be detected by any test suite. Higher order mutation testing has been coined as a promising solution for overcoming these limitations of FOM testing. In particular, finding strongly subsuming higher order mutants (SSHOMs), which are able to replace all of their constituent FOMs without scarifying test effectiveness while being able to reflect complex, real defects that require more than one change to correct them, is considered an important research challenge and is the focus of this work. The contribution of this article is a new, extended classification of higher order mutants (HOMs) to cover all cases of generated HOMs. Fitness functions and empirical comparison of four different multiobjective optimization algorithms are used to generate and evaluate HOMs as well as search for valuable high-quality and reasonable HOMs (strongly subsuming and coupled HOMs) and ten other types of HOMs. The main goal of this study is to assert the effect of applying multiobjective optimization algorithms in the area of higher order-mutation testing, while asserting the correctness of the proposed HOMs classification, objectives, and fitness functions. Our experimental results show that the total number of generated HOMs is smaller (about 70%) in comparison to FOMs, while the mean ratio of reasonable HOMs (subsuming HOMs) to all found HOMs is over 56%, and the mean ratio of high-quality and reasonable HOMs (strongly subsuming and coupled HOMs) to all found reasonable HOMs (subsuming HOMs) is fairly high (around 8.74%). Quang-Vu Nguyen 0001, Lech Madeyski |
Cybern. Syst. | 2 |
| 2015 | Which process metrics can significantly improve defect prediction models? An empirical studyabstractThe knowledge about the software metrics which serve as defect indicators is vital for the efficient allocation of resources for quality assurance. It is the process metrics, although sometimes difficult to collect, which have recently become popular with regard to defect prediction. However, in order to identify rightly the process metrics which are actually worth collecting, we need the evidence validating their ability to improve the product metric-based defect prediction models. This paper presents an empirical evaluation in which several process metrics were investigated in order to identify the ones which significantly improve the defect prediction models based on product metrics. Data from a wide range of software projects (both, industrial and open source) were collected. The predictions of the models that use only product metrics (simple models) were compared with the predictions of the models which used product metrics, as well as one of the process metrics under scrutiny (advanced models). To decide whether the improvements were significant or not, statistical tests were performed and effect sizes were calculated. The advanced defect prediction models trained on a data set containing product metrics and additionally Number of Distinct Committers (NDC) were significantly better than the simple models without NDC, while the effect size was medium and the probability of superiority (PS) of the advanced models over simple ones was high ( $$p=.016$$ , $$r=-.29$$ , $$\hbox {PS}=.76$$ ), which is a substantial finding useful in defect prediction. A similar result with slightly smaller PS was achieved by the advanced models trained on a data set containing product metrics and additionally all of the investigated process metrics ( $$p=.038$$ , $$r=-.29$$ , $$\hbox {PS}=.68$$ ). The advanced models trained on a data set containing product metrics and additionally Number of Modified Lines (NML) were significantly better than the simple models without NML, but the effect size was small ( $$p=.038$$ , $$r=.06$$ ). Hence, it is reasonable to recommend the NDC process metric in building the defect prediction models. Lech Madeyski, Marian Jureczko |
Softw. Qual. J. | 1 |
| 2014 | Overcoming the Equivalent Mutant Problem: A Systematic Literature Review and a Comparative Experiment of Second Order MutationabstractContext. The equivalent mutant problem (EMP) is one of the crucial problems in mutation testing widely studied over decades. Objectives. The objectives are: to present a systematic literature review (SLR) in the field of EMP; to identify, classify and improve the existing, or implement new, methods which try to overcome EMP and evaluate them. Method. We performed SLR based on the search of digital libraries. We implemented four second order mutation (SOM) strategies, in addition to first order mutation (FOM), and compared them from different perspectives. Results. Our SLR identified 17 relevant techniques (in 22 articles) and three categories of techniques: detecting (DEM); suggesting (SEM); and avoiding equivalent mutant generation (AEMG). The experiment indicated that SOM in general and JudyDiffOp strategy in particular provide the best results in the following areas: total number of mutants generated; the association between the type of mutation strategy and whether the generated mutants were equivalent or not; the number of not killed mutants; mutation testing time; time needed for manual classification. Conclusions . The results in the DEM category are still far from perfect. Thus, the SEM and AEMG categories have been developed. The JudyDiffOp algorithm achieved good results in many areas. Lech Madeyski, Wojciech Orzeszyna, Richard Torkar, Mariusz Jozala |
IEEE Trans. Software Eng. | 1 |
| 2013 | Continuous Test-Driven Development - A Novel Agile Software Development Practice and Supporting ToolabstractContinuous testing is a technique in modern software development in which the source code is constantly unit tested in the background and there is no need for the developer to perform the tests manually. We propose an extension to this technique that combines it with well-established software engineering practice called TestDriven Development (TDD). In our practice, that we called Continuous Test-Driven Development (CTDD), software developer writes the tests first and is not forced to perform them manually. We hope to reduce the time waste resulting from manual test execution in highly test driven development scenario. In this article we describe the CTDD practice and the tool that we intend to use to support and evaluate the CTDD practice in a real world software development project. Lech Madeyski, Marcin Kawalerowicz |
ENASE | 1 |
| 2010 | The impact of Test-First programming on branch coverage and mutation score indicator of unit tests: An experiment
Lech Madeyski |
Inf. Softw. Technol. | 1 |
| 2007 | The Impact of Test-Driven Development on Software Development Productivity - An Empirical Study
Lech Madeyski, Lukasz Szala |
EuroSPI | 1 |
| 2007 | On the Effects of Pair Programming on Thoroughness and Fault-Finding Effectiveness of Unit Tests
Lech Madeyski |
PROFES | 1 |
| 2007 | Empirical Evidence Principle and Joint Engagement Practice to Introduce XP
Lech Madeyski, Wojciech Biela |
XP | 1 |
| 2006 | The Impact of Pair Programming and Test-Driven Development on Package Dependencies in Object-Oriented Design - An Experiment
Lech Madeyski |
PROFES | 1 |
| 2006 | Is External Code Quality Correlated with Programming Experience or Feelgood Factor?
Lech Madeyski |
XP | 1 |