Sandro Morasca

dblp:m/SandroMorasca · DBLP profile ↗
← Back
78ranked-venue papers
22as first author
8since 2021 · last 2025
0000-0003-4598-7024ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 70 · 20 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorArtificial intelligence and machine learning · 4 · 3 first-authorSystems, architecture and hardware · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3Security and privacy · 1
YearPublicationVenuePosition
2025 Critical Considerations on Effort-aware Software Defect Prediction Metrics
abstract
Background. Effort-aware metrics (EAMs) are widely used to evaluate the effectiveness of software defect prediction models, while accounting for the effort needed to analyze the software modules that are estimated defective. The usual underlying assumption is that this effort is proportional to the modules’ size measured in LOC. However, the research on module analysis (including code understanding, inspection, testing, etc.) suggests that module analysis effort may be better correlated to code attributes other than size. Aim. We investigate whether assuming that module analysis effort is proportional to other code metrics than LOC leads to different evaluations. Method. We show mathematically that the choice of the code measure used as the module effort driver crucially influences the resulting evaluations. To illustrate the practical consequences of this, we carried out a demonstrative empirical study, in which the same model was evaluated via EAMs, assuming that effort is proportional to either McCabe’s complexity or LOC. Results. The empirical study showed that EAMs depend on the underlying effort model, and can give quite different indications when effort is modeled differently. It is also apparent that the extent of these differences varies widely. Conclusions. Researchers and practitioners should be aware that the reliability of the indications provided by EAMs depend on the nature of the underlying effort model. The EAMs used until now appear to be actually size-aware, rather than effort-aware: when analysis effort does not depend on size, these EAMs can be misleading.
Luigi Lavazza, Gabriele Rotoloni, Sandro Morasca
EASE3
2025 A Replicated Study on Factors Affecting Software Understandability
Georgia M. Kapitsaki, Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
ENASE3
2025 Software Defect Prediction evaluation: New metrics based on the ROC curve
abstract
Context: ROC (Receiver Operating Characteristic) curves are widely used to represent how well fault-proneness models (e.g., probability models) classify software modules as faulty or non-faulty. AUC , the Area Under the ROC Curve, is usually used to quantify the overall discriminating power of a fault-proneness model. Alternative indicators proposed, e.g., RRA (Ratio of Relevant Areas), consider the area under a portion of a ROC curve. Each point of a ROC curve represents a binary classifier, obtained by setting a specified threshold on the fault-proneness model. Several performance metrics (Precision, Recall, the F-score, etc.) are used to assess a binary classifier. Objectives: We investigate the relationships linking “under the ROC curve area” indicators such as AUC and RRA to performance metrics. Methods: We study these relationships analytically. We introduce iso-PM ROC curves, whose points have the same value P M ¯ for a given performance metric PM. When evaluating a ROC curve, we identify the iso-PM curve with the same value of AUC or RRA . Its P M ¯ can be seen as a property of the ROC curve and fault-proneness model under evaluation. Results: There is an S-shaped relationship between P M ¯ and AUC for performance metrics that do not depend on the proportion ρ of faulty modules, i.e., dataset balancedness. ϕ (Matthews Correlation Coefficient) depends on ρ : with very imbalanced datasets, AUC appears over-optimistic and ϕ over-pessimistic. RRA defines the region of interest in terms of ρ , so all performance metrics depend on ρ . RRA is related to performance metrics via S-shaped curves. Conclusion: Our proposal helps gain a better quantitative understanding of the goodness of a ROC curve, especially in practically relevant regions of interest. Also, showing a ROC curve and iso-PM curves provides an intuitive perception of the goodness of a fault-proneness model.
Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
Inf. Softw. Technol.2
2023 On the Reliability of the Area Under the ROC Curve in Empirical Software Engineering
abstract
Binary classifiers are commonly used in software engineering research to estimate several software qualities, e.g., defectiveness or vulnerability. Thus, it is important to adequately evaluate how well binary classifiers perform, before they are used in practice. The Area Under the Curve (AUC) of Receiver Operating Characteristic curves has often been used to this end. However, AUC has been the target of some criticisms, so it is necessary to evaluate under what conditions and to what extent AUC can be a reliable performance metric.
Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
EASE2
2023 An Experience in the Evaluation of Fault Prediction
Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
PROFES (1)2
2023 An empirical study on software understandability and its dependence on code characteristics
abstract
Abstract Context Insufficient code understandability makes software difficult to inspect and maintain and is a primary cause of software development cost. Several source code measures may be used to identify difficult-to-understand code, including well-known ones such as Lines of Code and McCabe’s Cyclomatic Complexity, and novel ones, such as Cognitive Complexity. Objective We investigate whether and to what extent source code measures, individually or together, are correlated with code understandability. Method We carried out an empirical study with students who were asked to carry out realistic maintenance tasks on methods from real-life Open Source Software projects. We collected several data items, including the time needed to correctly complete the maintenance tasks, which we used to quantify method understandability. We investigated the presence of correlations between the collected code measures and code understandability by using several Machine Learning techniques. Results We obtained models of code understandability using one or two code measures. However, the obtained models are not very accurate, the average prediction error being around 30%. Conclusions Based on our empirical study, it does not appear possible to build an understandability model based on structural code measures alone. Specifically, even the newly introduced Cognitive Complexity measure does not seem able to fulfill the promise of providing substantial improvements over existing measures, at least as far as code understandability prediction is concerned. It seems that, to obtain models of code understandability of acceptable accuracy, process measures should be used, possibly together with new source code measures that are better related to code understandability.
Luigi Lavazza, Sandro Morasca, Marco Gatto
Empir. Softw. Eng.2
2023 An empirical evaluation of the "Cognitive Complexity" measure as a predictor of code understandability
Luigi Lavazza, Abedallah Zaid Abualkishik, Sandro Morasca
J. Syst. Softw.4
2022 Comparing ϕ and the F-measure as performance metrics for software-related classifications
abstract
Abstract Context The F-measure has been widely used as a performance metric when selecting binary classifiers for prediction, but it has also been widely criticized, especially given the availability of alternatives such as ϕ (also known as Matthews Correlation Coefficient). Objectives Our goals are to (1) investigate possible issues related to the F-measure in depth and show how ϕ can address them, and (2) explore the relationships between the F-measure and ϕ . Method Based on the definitions of ϕ and the F-measure, we derive a few mathematical properties of these two performance metrics and of the relationships between them. To demonstrate the practical effects of these mathematical properties, we illustrate the outcomes of an empirical study involving 70 Empirical Software Engineering datasets and 837 classifiers. Results We show that ϕ can be defined as a function of Precision and Recall , which are the only two performance metrics used to define the F-measure, and the rate of actually positive software modules in a dataset. Also, ϕ can be expressed as a function of the F-measure and the rates of actual and estimated positive software modules. We derive the minimum and maximum value of ϕ for any given value of the F-measure, and the conditions under which both the F-measure and ϕ rank two classifiers in the same order. Conclusions Our results show that ϕ is a sensible and useful metric for assessing the performance of binary classifiers. We also recommend that the F-measure should not be used by itself to assess the performance of a classifier, but that the rate of positives should always be specified as well, at least to assess if and to what extent a classifier performs better than random classification. The mathematical relationships described here can also be used to re-interpret the conclusions of previously published papers that relied mainly on the F-measure as a performance metric.
Luigi Lavazza, Sandro Morasca
Empir. Softw. Eng.2
2020 Open Source Software Evaluation, Selection, and Adoption: a Systematic Literature Review
abstract
Background. Open Source Software (OSS) is experiencing an increasing popularity both in industry and in academia. Aim. We investigated models for the selection, evaluation, and adoption of OSS, focusing on factors that affect most the evaluation of OSS. Method. We conducted a Systematic Literature Review of 262 studies published until the end of 2019, to understand whether OSS selection is still an interesting topic for researchers, and which factors are considered by stakeholders and are assessed by the available models. Result. We selected 60 primary studies: 20 surveys and 5 lessons learned studies elicited the motivations for OSS adoption; 35 papers proposed several OSS evaluation models focusing on different technical aspects. This Systematic Literature Review provides an overview of the available OSS evaluation methods, highlighting their limits and strengths, based on the wide range of technicalities and aspects explored by the selected primary studies. Conclusion. OSS producers can benefit from our results by checking if they are providing all the information commonly required by potential adopters. Users can learn how models work and which models cover the relevant characteristics of OSS they are most interested in.
Valentina Lenarduzzi, Davide Taibi 0001, Davide Tosi, Luigi Lavazza, Sandro Morasca
SEAA5
2020 An Empirical Study of Thresholds for Code Measures
abstract
Background. Practical use of software code measures for faultiness estimation often requires setting thresholds on measures, to separate software modules that are likely to be faulty from modules that are likely to be non-faulty. Several threshold proposals exist in the literature. However, different proposals may recommend different threshold values for the same measure, so practitioners may be unsure which threshold value they should use.Objective. Our goal is to investigate whether it is possible to define a single threshold for a code measure, or at least a small range within which to select a threshold.Method. We carried out an empirical study based on two collections of datasets available in the SEACRAFT repository. For each dataset, we built all statistically significant univariate Binary Logistic Regression fault-proneness models, using a set of the most commonly used code measures as independent variables. We then derived thresholds for single software measures by setting an acceptability threshold to fault-proneness. We then checked whether the distribution of the thresholds obtained for the same measure is concentrated enough that it is possible to take a value of the code measure as a “universal” threshold for it. We repeated the same method with bivariate Binary Logistic Regression fault-proneness models.Results. The threshold distributions we obtained are quite dispersed, so it does not seem to be possible to define “universal” thresholds for the code measures we considered.Conclusions. According to the collected evidence, it appears hardly possible to effectively control software faultiness by recommending that the most commonly used code measures satisfy fixed thresholds.
Luigi Lavazza, Sandro Morasca
ISSRE2
2020 On the assessment of software defect prediction models via ROC curves
abstract
Abstract Software defect prediction models are classifiers often built by setting a threshold t on a defect proneness model, i.e., a scoring function. For instance, they classify a software module non-faulty if its defect proneness is below t and positive otherwise. Different values of t may lead to different defect prediction models, possibly with very different performance levels. Receiver Operating Characteristic (ROC) curves provide an overall assessment of a defect proneness model, by taking into account all possible values of t and thus all defect prediction models that can be built based on it. However, using a defect proneness model with a value of t is sensible only if the resulting defect prediction model has a performance that is at least as good as some minimal performance level that depends on practitioners’ and researchers’ goals and needs. We introduce a new approach and a new performance metric (the Ratio of Relevant Areas) for assessing a defect proneness model by taking into account only the parts of a ROC curve corresponding to values of t for which defect proneness models have higher performance than some reference value. We provide the practical motivations and theoretical underpinnings for our approach, by: 1) showing how it addresses the shortcomings of existing performance metrics like the Area Under the Curve and Gini’s coefficient; 2) deriving reference values based on random defect prediction policies, in addition to deterministic ones; 3) showing how the approach works with several performance metrics (e.g., Precision and Recall) and their combinations; 4) studying misclassification costs and providing a general upper bound for the cost related to the use of any defect proneness model; 5) showing the relationships between misclassification costs and performance metrics. We also carried out a comprehensive empirical study on real-life data from the SEACRAFT repository, to show the differences between our metric and the existing ones and how more reliable and less misleading our metric can be.
Sandro Morasca, Luigi Lavazza
Empir. Softw. Eng.1
2019 Dealing with Uncertainty in Binary Logistic Regression Fault-proneness Models
abstract
Background Binary Logistic Regression is widely used in Empirical Software Engineering to build estimation models, e.g., fault-proneness models, which estimate the probability that a given module is faulty, based on some measures of the module. Fault-proneness models are then used to build faultiness model, i.e., models that estimate whether a given module is faulty or non-faulty.
Luigi Lavazza, Sandro Morasca
EASE2
2019 Comparing the Effectiveness of Using Design and Code Measures in Software Faultiness Estimation
abstract
Background. Early identification of software modules that are likely to be faulty helps practitioners take timely actions to improve these modules' quality and reduce development costs in the remainder of the development process. To this end, module faultiness estimation models can be built at any point during development by using measures collected up to that time. Models available in later phases are expected to be more accurate than those available in earlier phases. However, waiting until late in the development process may reduce the impact of the effectiveness and efficacy of any software quality improvement actions and increase their cost.
Sandro Morasca, Luigi Lavazza
EASE1
2019 Empirical evaluation and proposals for bands-based COSMIC early estimation methods
Luigi Lavazza, Sandro Morasca
Inf. Softw. Technol.2
2018 Technical debt as an external software attribute
abstract
Background: Technical debt is currently receiving increasing attention from practitioners and researchers. Several metaphors, concepts, and indications concerning technical debt have been introduced, but no agreement exists about a solid definition of technical debt.
Luigi Lavazza, Sandro Morasca, Davide Tosi
TechDebt@ICSE2
2018 Investigating the impact of fault data completeness over time on predicting class fault-proneness
Jehad Al Dallal, Sandro Morasca
Inf. Softw. Technol.2
2017 On the Evaluation of Effort Estimation Models
abstract
Background. Using accurate effort estimation models can help software companies plan, monitor, and control their development process and development costs. It is therefore important to define sound accuracy indicators that allow practitioners and researchers to assess and rank different effort estimation models so that practitioners can select the most accurate, and therefore useful one. Several accuracy indicators exist, with different advantages and disadvantages.
Luigi Lavazza, Sandro Morasca
EASE2
2017 Risk-averse slope-based thresholds: Definition and empirical evaluation
Sandro Morasca, Luigi Lavazza
Inf. Softw. Technol.1
2016 Slope-based fault-proneness thresholds for software engineering measures
abstract
Background. Practical use of a measure X for an internal attribute (e.g., size, structural complexity, cohesion, coupling) of a software module often requires setting a threshold on X, to make decisions as to which software modules may be estimated to be potentially faulty. To keep quality under control, practitioners may want to set a threshold on X to identify "early symptoms" of possible faultiness of a module, which should be closely monitored and possibly modified.
Sandro Morasca, Luigi Lavazza
EASE1
2016 Identifying Thresholds for Software Faultiness via Optimistic and Pessimistic Estimations
abstract
Background. When estimating whether a software module is faulty based on the value of a measure X for a software internal attribute (e.g., size, structural complexity, cohesion, coupling), it is sensible to set a threshold on fault-proneness first and then induce a threshold on X by using a fault-proneness model where X plays the role of independent variable. However, some modules cannot be estimated as either faulty or non-faulty with confidence: they belong to a "grey zone" and estimating them as either would be quite aleatory and may result in several erroneous decisions.
Luigi Lavazza, Sandro Morasca
ESEM2
2016 An Empirical Evaluation of Two COSMIC Early Estimation Methods
abstract
Background. Under specific circumstances -especially in the early phases of software development projects-a thorough application of the COSMIC method may require more time and effort than available. Thus, early approximate estimation methods have been proposed for estimating the functional size of a given application, instead of properly measuring it. Objective. This paper aims at empirically evaluating the accuracy of two COSMIC early size estimation methods. The goal is to provide practitioners with some empirical evidence on the accuracy that can be expected from these methods. Method. We evaluated the Average Functional Process and the Equal Size Bands methods by applying them to a set of applications that were previously measured according to the standard COSMIC method, and for which the data necessary to perform estimations were readily available. The application conditions and performance of the Equal Size Bands method were also evaluated from a theoretical point of view. Results. Our analyses show that in a few cases the Average Functional Process method features estimation errors that are too large to be acceptable, while on average it provides reasonable estimates. On the contrary, the Equal Size Bands method can provide quite accurate estimates, but only if the human measurer is sufficiently good at classifying each functional process in the correct band. From a theoretical point of view, it is shown that-perhaps counterintuitively- the Average Functional Process and the Equal Size Bands methods provide essentially equivalent estimates when the distribution of Functional Processes across the bands is the same in the historical datasets and in the new software to be estimated. When such distributions are quite different, experimental results show that the Equal Size Bands method performs much better than the Average Functional Process method. Conclusions. Our results show that in a few cases the Average Functional Process method fails to provide acceptably small estimation errors. On the contrary, the Equal Size Bands method is sufficiently accurate to provide good size estimates. However, organizations that plan to use it need to properly train measurers that are able to identify the correct size band in which every functional process belongs.
Luigi Lavazza, Sandro Morasca
IWSM-Mensura2
2015 Classifying faulty modules with an extension of the H-index
abstract
Background. Fault-proneness estimation models provide an estimate of the probability that a software module is faulty. These models can be used to classify modules as faulty or non-faulty, by using a fault-proneness threshold: modules whose fault-proneness exceeds the threshold are classified as faulty and the others as non-faulty. However, the selection of the threshold value is to some extent subjective, and different threshold values may lead to very different results in terms of classification accuracy. Objective. We introduce and empirically validate a new approach to setting thresholds, based on an extension of the H-index defined in Bibliometrics, called the Fault-proneness H-Index. We define and use this extension to identify the most fault-prone modules, which are candidates for intensive Verification & Validation activities. Method. We carried out the empirical validation on two data sets with different faultiness characteristics hosted on the PROMISE repository, by using T-times K-fold cross validation. We computed Precision, Recall, the F - measure, and a weighted version of the F - measure for the results obtained with our approach and compared them with the values obtained with other approaches based on several thresholds. Results. In the empirical validation, our approach provides better classification results than those based on most other thresholds, according to some classification accuracy indicators, in a statistically significant way. Conclusions. Our approach seems to be able to contribute to accurately classifying modules as faulty or non-faulty.
Sandro Morasca
ISSRE1
2015 Supporting the semi-automatic semantic annotation of web services: A systematic literature review
Davide Tosi, Sandro Morasca
Inf. Softw. Technol.2
2014 Using logistic regression to estimate the number of faulty software modules
abstract
Background. The evaluation of the accuracy of an estimation model for software fault-proneness is carried out by using the model with data collected on a set of software modules and classifying the modules in the set as either estimated faulty or estimated non-faulty. This classification usually involves setting a fault-proneness threshold: software modules whose fault-proneness is above that threshold are classified as estimated faulty and the others as estimated non-faulty. The selection of the threshold value is to some extent subjective and arbitrary, and different threshold values may lead to very different results in terms of classification accuracy.
Sandro Morasca
EASE1
2014 Predicting object-oriented class reuse-proneness using internal quality attributes
Jehad Al Dallal, Sandro Morasca
Empir. Softw. Eng.2
2014 Model-based early and rapid estimation of COSMIC functional size - An experimental evaluation
Vieri Del Bianco, Luigi Lavazza, Sandro Morasca, Abedallah Zaid Abualkishik
Inf. Softw. Technol.4
2013 Towards a simplified definition of Function Points
abstract
The measurement of Function Points is based on Base Functional Components. The process of identifying and weighting Base Functional Components is hardly automatable, due to the informality of both the Function Point method and the requirements documents being measured. So, Function Point measurement generally requires a lengthy and costly process. We investigate whether it is possible to take into account only subsets of Base Functional Components so as to obtain functional size measures that simplify Function Points with the same effort estimation accuracy as the original Function Points measure. Simplifying the definition of Function Points would imply a reduction of measurement costs and may help spread the adoption of this type of measurement practices. Specifically, we empirically investigate the following issues: whether available data provide evidence that simplified software functionality measures can be defined in a way that is consistent with Function Point Analysis; whether simplified functional size measures by themselves can be used without any appreciable loss in software development effort prediction accuracy; whether simplified functional size measures can be used as software development effort predictors in models that also use other software requirements measures. We analyze the relationships between Function Points and their Base Functional Components. We also analyze the relationships between Base Functional Components and development effort. Finally, we built effort prediction models that contain both the simplified functional measures and additional requirements measures. Significant statistical models correlate Function Points with Base Functional Components. Basic Functional Components can be used to build models of effort that are equivalent, in terms of accuracy, to those based on Function Points. Finally, simplified Function Points measures can be used as software development effort predictors in models that also use other requirements measures. The definition and measurement processes of Function Points can be dramatically simplified by taking into account a subset of the Base Functional Components used in the original definition of the measure, thus allowing for substantial savings in measurement effort, without sacrificing the accuracy of software development effort estimates.
Luigi Lavazza, Sandro Morasca, Gabriela Robiolo
Inf. Softw. Technol.2
2013 A systematic review on the functional testing of semantic web services
Abbas Tahir, Davide Tosi, Sandro Morasca
J. Syst. Softw.3
2012 Software effort estimation with a generalized robust linear regression technique
abstract
Background. Outliers and corrupted data points may unduly bias software development effort estimation models. However, given the usually limited size of software engineering data sets, removing too many data points may seriously reduce the power of the statistical tests used and the likelihood of statistically significant result. Also, statistical techniques are typically based on assumptions that are either believed to be true a priori or, at best, checked via statistical tests, without ever achieving 100% certainty on their truthfulness. Estimation models based on less strict assumptions have broader applicability and lower risks of drawing unwarranted conclusions. Aim. We investigate the usefulness of Robust Regression when building effort estimation models, by varying the degree of robustness and, thus, the number of data points that are excluded from the data analysis as outliers. Method. We have used Least Quantile of Squares (LQS) Robust Regression, a generalization of the Least Median of Squares (LMS). LMS builds a regression line by minimizing the median squared residual. LQS minimizes the order statistic of square residuals corresponding to any specified quantile, and not just the median, which is the order statistic corresponding to the 50% quantile. We have extended a statistical significance test for univariate LQS regression models. We have also built a weighted model, obtained from statistically significant LQS models, where each LQS model contributes proportionally to the quantile used. Results. We have applied LQS Linear Regression to estimate development effort on four projects from the PROMISE data set and obtained valid and significant univariate models. Conclusions. LQS may provide a valid alternative to LMS and Ordinary Least Square regressions to build estimation models when (1) balancing the need for excluding outliers and keeping enough data points to build statistically significant models and (2) using less strict assumptions underlying the regression technique.
Luigi Lavazza, Sandro Morasca
EASE2
2012 On the definition of dynamic software measures
abstract
The quantification of several software attributes (e.g., size, complexity, cohesion, coupling) is usually carried out in a static fashion, and several hundreds of measures have been defined to this end. However, static measurement may only be an approximation for the measurement of these attributes during software use. The paper proposes a theoretical framework based on Axiomatic Approaches for the definition of sensible dynamic software measures that can dynamically capture these attributes. Dynamic measures based on this framework are defined for dynamically quantifying size and coupling. In this paper, we also compare dynamic measures of size and coupling against well-known static measures by correlating them with fault-pronenesses of four case studies.
Davide Tosi, Luigi Lavazza, Sandro Morasca, Davide Taibi 0001
ESEM3
2012 A Proposal for Simplified Model-Based Cost Estimation Models
Vieri Del Bianco, Luigi Lavazza, Sandro Morasca
PROFES3
2011 Software Measures for Business Processes
Giuseppe Pozzi, Sandro Morasca, Alessio Antonini, Alexandre Mello Ferreira
ADBIS (2)2
2011 A probability-based approach to modeling the risk of unauthorized propagation of information in on-line social networks
abstract
The unauthorized propagation of information is an important problem in the Internet, especially because of the increasing popularity of On-line Social Networks. To address this issue, many access control mechanisms have been proposed so far, but there is still a lack of techniques to evaluate the risk of unauthorized flow of information within social networks. This paper introduces a probability-based approach to modeling the likelihood that information propagates from one social network user to users who are not authorized to access it. The approach is demonstrated via an example, to show how it can be applied in practical cases.
Barbara Carminati, Elena Ferrari 0001, Sandro Morasca, Davide Taibi 0001
CODASPY3
2011 OP2A - Assessing the Quality of the Portal of Open Source Software Products
Gabriele Basilico, Luigi Lavazza, Sandro Morasca, Davide Taibi 0001, Davide Tosi
WEBIST3
2011 Convertibility of Function Points into COSMIC Function Points: A study using Piecewise Linear Regression
Luigi Lavazza, Sandro Morasca
Inf. Softw. Technol.2
2010 Predicting OSS trustworthiness on the basis of elementary code assessment
abstract
Background. Open Source Software (OSS) provides increasingly serious and viable alternatives to traditional closed source software. The number of OSS users is continuously growing, as is the number of potential users that are interested in evaluating the quality of OSS. The latter would greatly benefit from simple methods for evaluating the trustworthiness of OSS.
Luigi Lavazza, Sandro Morasca, Davide Taibi 0001, Davide Tosi
ESEM2
2010 Applying SCRUM in an OSS Development Process: An Empirical Evaluation
Luigi Lavazza, Sandro Morasca, Davide Taibi 0001, Davide Tosi
XP2
2010 A checklist for integrating student empirical studies with research and teaching goals
Jeffrey C. Carver, Letizia Jaccheri, Sandro Morasca, Forrest Shull
Empir. Softw. Eng.3
2009 A probability-based approach for measuring external attributes of software artifacts
abstract
The quantification of so-called external software attributes, which are the product qualities with real relevance for developers and users, has often been problematic. This paper introduces a proposal for quantifying external software attributes in a unified way. The basic idea is that external software attributes can be quantified by means of probabilities. As a consequence, external software attributes can be estimated via probabilistic models, and not directly measured via software measures. This paper discusses the reasons underlying the proposals and shows the pitfalls related to using measures for external software attributes. We also show that the theoretical bases for our approach can be found in so-called ldquoprobability representations,rdquo a part of Measurement Theory that has not yet been used in Software Engineering Measurement. By taking the definition and estimation of reliability as reference, we show that other external software attributes can be defined and modeled by a probability-based approach.
Sandro Morasca
ESEM1
2009 3rd International Workshop on Designing Empirical Studies: Assessing the Effectiveness of Agile Methods (IWDES 2009)
Massimiliano Di Penta, Sandro Morasca, Alberto Sillitti
XP2
2009 Assessing the understandability of UML statechart diagrams with composite states - A family of empirical studies
José A. Cruz-Lemus, Marcela Genero, M. Esperanza Manso, Sandro Morasca, Mario Piattini
Empir. Softw. Eng.4
2008 Refining the axiomatic definition of internal software attributes
abstract
Several internal software attributes, like size, structural complexity, cohesion, coupling, have been introduced and used to reason about software engineering artifacts, and many measures have been proposed for them. Internal software attributes are important because they are believed to be related to quantities of industrial interest, like the number of defects or the development effort. However, the definition of internal software attributes still needs to be made more precise and formal, so measures can be defined that really quantify the attributes they purport to measure. In this paper, we extend, simplify, and refine an existing axiomatic approach that characterizes each internal attribute rigorously via a different set of axioms. This paper makes three specific contributions. First, the new proposal captures a larger set of aspects of software artifacts that may be relevant for internal software attributes than the original proposal did. Second, we identify the basic, foundational sets of axioms for each internal attribute studied, from which the other properties of the attribute can be derived, so the intrinsic properties of the attribute and their implications can be understood. Third, we investigate some relevant relationships among internal software attributes, so their similarities and differences, which are sometimes not well identified, are made more explicit.
Sandro Morasca
ESEM1
2008 Subjective Assessment of the Mutual Influence of ISO 9126 Software Qualities: an Empirical Study
Sandro Morasca
SEKE1
2007 Three empirical studies on estimating the design effort of Web applications
abstract
Our research focuses on the effort needed for designing modern Web applications. The design effort is an important part of the total development effort, since the implementation can be partially automated by tools. We carried out three empirical studies with students of advanced university classes enrolled in engineering and communication sciences curricula. The empirical studies are based on the use of W2000, a special-purpose design notation for the design of Web applications, but the hypotheses and results may apply to a wider class of modeling notations (e.g., OOHDM, WebML, or UWE). We started by investigating the relative importance of each design activity. We then assessed the accuracy of a priori design effort predictions and the influence of a few process-related factors on the effort needed for each design activity. We also analyzed the impact of attributes like the size and complexity of W2000 design artifacts on the total effort needed to design the user experience of web applications. In addition, we carried out a finer-grain analysis, by studying which of these attributes impact the effort devoted to the steps of the design phase that are followed when using W2000.
Luciano Baresi, Sandro Morasca
ACM Trans. Softw. Eng. Methodol.2
2006 On the Assessment of the Mean Failure Frequency of Software in Late Testing
abstract
We propose an approach for assessing the mean failure frequency of a program, based on the statistical test of hypotheses. The approach can be used to establish stopping rules and evaluate the quality of a program based on its mean failure frequency during the late testing phases. Our proposal shows how to set and satisfy conservative bounds for the minimum number of test executions that are needed to achieve a target mean failure frequency with a specified level of statistical significance, based on the quality goal of testing and the specific test execution profile chosen. We relax a few assumptions of the literature, so our approach can be used in a larger set of real-life cases.
Sandro Morasca
ICSEA1
2005 Towards Model-Driven Testing of a Web Application Generator
Luciano Baresi, Piero Fraternali, Massimo Tisi, Sandro Morasca
ICWE4
2004 On the analytical comparison of testing techniques
abstract
We introduce necessary and sufficient conditions for comparing the expected values of the number of failures caused by applications of software testing techniques. Our conditions are based only on the knowledge of a total or even a hierarchical order among the failure rates of the subdomains of a program's input domain. We also prove conditions for comparing the probability of causing at least one failure in three important special cases.
Sandro Morasca, Stefano Serra-Capizzano
ISSTA1
2003 Foundations of a Weak Measurement-Theoretic Approach to Software Measurement
Sandro Morasca
FASE1
2003 A Bayesian Approach to Software Testing Evaluation
Sandro Morasca
SEKE1
2003 Towards Industrially Relevant Fault-Proneness Models
abstract
Estimating software fault-proneness early, i.e., predicting the probability of software modules to be faulty, can help in reducing costs and increasing effectiveness of software analysis and testing. The many available static metrics provide important information, but none of them can be deterministically related to software fault-proneness. Fault-proneness models seem to be an interesting alternative, but the work on these is still biased by lack of experimental validation. This paper discusses barriers and problems in using software fault-proneness in industrial environments, proposes a method for building software fault-proneness models based on logistic regression and cross-validation that meets industrial needs, and provides some experimental evidence of the validity of the proposed approach.
Giovanni Denaro, Mauro Pezzè, Sandro Morasca
Int. J. Softw. Eng. Knowl. Eng.3
2002 Deriving models of software fault-proneness
abstract
The effectiveness of the software testing process is a key issue for meeting the increasing demand of quality without augmenting the overall costs of software development. The estimation of software fault-proneness is important for assessing costs and quality and thus better planning and tuning the testing process. Unfortunately, no general techniques are available for estimating software fault-proneness and the distribution of faults to identify the correct level of test for the required quality. Although software complexity and testing thoroughness are intuitively related to the costs of quality assurance and the quality of the final product, single software metrics and coverage criteria provide limited help in planning the testing process and assuring the required quality.By using logistic regression, this paper shows how models can be built that relate software measures and software fault-proneness for classes of homogeneous software products. It also proposes the use of cross-validation for selecting valid models even for small data sets.The early results show that it is possible to build statistical models based on historical data for estimating fault-proneness of software modules before testing, and thus better planning and monitoring the testing activities.
Giovanni Denaro, Sandro Morasca, Mauro Pezzè
SEKE2
2002 A proposal for using continuous attributes in classification trees
abstract
Classification trees have been successfully used in several application fields. However, continuous attributes cannot be used directly when building classification trees, but they must be first discretized with clustering techniques, which require some degree of subjectivity. We propose an approach to build classification trees that does not require the discretization of the continuous attributes. The approach is an extension of existing methods for building classification trees and is based on the information gain yielded by discrete and continuous attributes. Data from a software development case study are analyzed with both the proposed approach and C4.5 to show the approach's applicability and benefits over C4.5.
Sandro Morasca
SEKE1
2002 An Empirical Study on the Design Effort of Web Applications
abstract
We study the effort needed for designing Web applications from an empirical point of view. The design phase forms an important part of the overall effort needed to develop a Web application, since the use of tools can help automate the implementation phase. We carried out an empirical study with students of an advanced university class that used W2000 as a Web application design technique. Our first goal was to compare the relative importance of each design activity. Second, we tried to assess the accuracy of a priori design effort predictions and the influence of factors on the effort needed for each design activity. Third, we also studied the quality of the designs obtained.
Luciano Baresi, Sandro Morasca, Paolo Paolini
WISE2
2002 An Operational Process for Goal-Driven Definition of Measures
abstract
We propose an approach (GOM/MEDEA) for defining measures of product attributes in software engineering. The approach is driven by the experimental goals of measurement, expressed via the GQM paradigm, and a set of empirical hypotheses. To make the empirical hypotheses quantitatively verifiable, GQM/MEDEA supports the definition of theoretically valid measures for the attributes of interest based on their expected mathematical properties. The empirical hypotheses are subject to experimental verification. This approach integrates several research contributions from the literature into a consistent, practical, and rigorous approach.
Lionel C. Briand, Sandro Morasca, Victor R. Basili
IEEE Trans. Software Eng.2
2001 An Empirical Study of Software Productivity
abstract
We studied productivity in a real-life environment in the Italian public administration by applying the goal/question/metrics paradigm to define a productivity related measurement goal and derive measures that were deemed relevant to reach the stated goal. Productivity was studied from both a functional and a product size perspectives. Our study has highlighted a few factors that are related to either aspect of productivity. The results may provide software managers with support for evaluating and improving software processes, so they can make decisions based on more quantitative information.
Sandro Morasca, Giuliano Russo
COMPSAC1
2000 A Case Study on Applying a Tool for Automated System Analysis Based on Modular Specifications Written in TRIO
Sandro Morasca, Angelo Morzenti, Pierluigi San Pietro
Autom. Softw. Eng.1
2000 A hybrid approach to analyze empirical software engineering data and its application to predict module fault-proneness in maintenance
Sandro Morasca, Günther Ruhe
J. Syst. Softw.1
2000 Generation of Execution Sequences for Modular Time Critical Systems
abstract
We define methods for generating execution sequences for time-critical systems based on their modularized formal specification. An execution sequence represents a behavior of a time critical system and can be used, before the final system is built, to validate the system specification against the user requirements (specification validation) and, after the final system is built, to verify whether the implementation satisfies the specification (functional testing). Our techniques generate execution sequences in the large, in that we focus on the connections among the abstract interfaces of the modules composing a modular specification. Execution sequences in the large are obtained by composing execution sequences in the small for the individual modules. We abstract from the specification languages used for the individual modules of the system, so our techniques can also be used when the modules composing the system are specified with different formalisms. We consider the cases in which connections give rise to either circular or noncircular dependencies among specification modules. We show that execution sequence generation can be carried out successfully under rather broad conditions and we define procedures for efficient construction of execution sequences. These procedures can be taken as the basis for the implementation of (semi)automated tools that provide substantial support to the activity of specification validation and functional testing for industrially-sized time critical systems.
Pierluigi San Pietro, Angelo Morzenti, Sandro Morasca
IEEE Trans. Software Eng.3
1999 Defining and Validating Measures for Object-Based High-Level Design
abstract
The availability of significant measures in the early phases of the software development life-cycle allows for better management of the later phases, and more effective quality assessment when quality can be more easily affected by preventive or corrective actions. We introduce and compare various high-level design measures for object-based software systems. The measures are derived based on an experimental goal, identifying fault-prone software parts, and several experimental hypotheses arising from the development of Ada systems for Flight Dynamics Software at the NASA Goddard Space Flight Center (NASA/GSFC). Specifically, we define a set of measures for cohesion and coupling, which satisfy a previously published set of mathematical properties that are necessary for any such measures to be valid. We then investigate the measures' relationship to fault-proneness on three large scale projects, to provide empirical support for their practical significance and usefulness.
Lionel C. Briand, Sandro Morasca, Victor R. Basili
IEEE Trans. Software Eng.2
1998 A Tool for Automated System Analysis based on Modular Specifications
abstract
An effective means for analyzing and reasoning on software systems is to use formal specifications to simulate their execution. The simulation traces can be used for specification testing and reused, later in the development process, for functional testing of the system. It is widely acknowledged that, to deal with the complexity of industrial-size systems, specifications must be structured into modules providing abstraction mechanisms and clear interfaces. In past work (D. Mandrioloi et al., 1995), we defined and implemented a method for simulating specifications written in the TRIO temporal logic language, and applied it to functional testing of time-critical industrial systems. In this paper, we report on a tool for analyzing TRIO specifications taking advantage of their modular structure, overcoming the well-known state-explosion problem and making the proposed method really scalable. We discuss the fundamental operations and the algorithms on which the tool is based. Then we illustrate its use in a realistic case study inspired by an industrial application. Finally, we comment on the overall results in terms of the usability of the tool and the effectiveness of the approach, and we suggest some future improvements.
Angelo Morzenti, Pierluigi San Pietro, Sandro Morasca
ASE3
1998 Applying GQM in an Industrial Software Factory
abstract
Goal/Question/Metric GQM) is a paradigm for the systematic definition, establishment, and exploitation of measurement programs supporting the quantitative evaluation of softare processes and products. Although GQM is a quite well-known method, detailed guidelines for establishing a GQM program in an industrial environment are still limited. Also, there are few reported experiences on the application of GQM to industrial cases. Finally, the technological support for GQM is still inadequate. This article describes the experience we have gained in applying GQM at Digital Laboratries in Italy. The procedures, experiences, and technology that have been employed in this study are largely reusable by other industrial organizations willing to introduce a GQM-based measurement program in their development environments.
Alfonso Fuggetta, Luigi Lavazza, Sandro Morasca, Stefano Cinti, Giandomenico Oldano, Elena Orazi
ACM Trans. Softw. Eng. Methodol.3
1997 Software Measurement and Formal Methods: A Case Study Centered on TRIO+ Specifications
abstract
Presents a case study where product measures are defined for a formal specification language (TRIO+) and are validated as quality indicators. To this end, defect and effort data were collected during the development of a monitoring and control system for a power plant. We show that some of the underlying hypotheses of these measures are supported bp empirical results and that several measures are significant early indicators of specification change and effort. From a more general perspective, this study exemplifies one important advantage of formal specifications: they are measurable and can thus be better controlled, assessed and managed than informal ones.
Lionel C. Briand, Sandro Morasca
ICFEM2
1997 Reply to ''Comments to the Paper: Briand, El Emam, Morasca: On the Application of Measurement Theory in Software Engineering
Lionel C. Briand, Khaled El Emam, Sandro Morasca
Empir. Softw. Eng.3
1997 Applying QIP/GQM in a Maintenance Project
Sandro Morasca
Empir. Softw. Eng.1
1997 Response to: Comments on "Property-Based Software Engineering Measurement: Refining the Additivity Properties"
abstract
S.196-197 : Lit.
Lionel C. Briand, Sandro Morasca, Victor R. Basili
IEEE Trans. Software Eng.2
1997 Comments on "Toward a Framework for Software Measurement Validation"
abstract
A view of software measurement that disagrees with the model presented by Kitchenham, Pfleeger, and Fenton (1995), is given. Whereas Kitchenham et al. argue that properties used to define measures should not constrain the scale type of measures, the authors contend that that is an inappropriate restriction. In addition, a misinterpretation of Weyuker's (1988) properties is noted.
Sandro Morasca, Lionel C. Briand, Victor R. Basili, Elaine J. Weyuker, Marvin V. Zelkowitz
IEEE Trans. Software Eng.1
1996 Generating Functional Test Cases in-the-large for Time-critical Systems from Logic-based Specifications
abstract
We address the problem of generating functional test cases for complex, highly structured time-critical systems starting from a modularized logic-based specification written in the TRIOR+ language, an object-oriented extension of the temporal logic TRIO.First, we present methods for producing test cases for a TRIO+ specification module, referring both to the internal, hidden, portion of the module and to its interface. Then, we discuss criteria to be used in the construction of test cases from a TRIO+ specification based on its composing modules and the connections among their interfaces. We formally define the notions related to test case derivation from TRIO+ modules and we introduce an executable language for describing a variety of strategies for constructing test cases for structured TRIO+ specifications starting from (parts of) the test cases of the composing modules. This language can be the basis for the implementation of an interactive tool for the semiautomatic construction of functional test cases from complex time-critical systems starting from their TRIO+ specification.
Sandro Morasca, Angelo Morzenti, Pierluigi San Pietro
ISSTA1
1996 On the application of measurement theory in software engineering
Lionel C. Briand, Khaled El Emam, Sandro Morasca
Empir. Softw. Eng.3
1996 Assessment of fault-detection processes: an approach based on reliability techniques
abstract
Two major factors influence the number of faults uncovered by a fault-detection process applied to a software artifact (e.g., specification, code): ability of the process to uncover faults, quality of the artifact (number of existing faults). These two factors must be assessed separately, so that one can: switch to a different process if the one being used is not effective enough, or stop the process if the number of remaining faults is acceptable. The fault-detection process assessment model can be applied to all sorts of artifacts produced during software development, and provides measures for both the 'effectiveness of a fault-detection process' and the 'number of existing faults in the artifact'. The model is valid even when there are zero defects in the artifact or the fault-detection process is intrinsically unable to uncover faults. More specifically, the times between fault discoveries are modeled via reliability-based techniques with an exponential distribution. The hazard rate is the product of 'effectiveness of the fault-detection process' and 'number of faults in the artifact'. Based on general hypotheses, the number of faults in an artifact follows a Poisson distribution. The unconditional distribution, whose parameters are estimated via maximum likelihood, is obtained.
Sandro Morasca
IEEE Trans. Reliab.1
1996 Property-Based Software Engineering Measurement
abstract
Little theory exists in the field of software system measurement. Concepts such as complexity, coupling, cohesion or even size are very often subject to interpretation and appear to have inconsistent definitions in the literature. As a consequence, there is little guidance provided to the analyst attempting to define proper measures for specific problems. Many controversies in the literature are simply misunderstandings and stem from the fact that some people talk about different measurement concepts under the same label (complexity is the most common case). There is a need to define unambiguously the most important measurement concepts used in the measurement of software products. One way of doing so is to define precisely what mathematical properties characterize these concepts, regardless of the specific software artifacts to which these concepts are applied. Such a mathematical framework could generate a consensus in the software engineering community and provide a means for better communication among researchers, better guidelines for analysts, and better evaluation methods for commercial static analyzers for practitioners. We propose a mathematical framework which is generic, because it is not specific to any particular software artifact, and rigorous, because it is based on precise mathematical concepts. We use this framework to propose definitions of several important measurement concepts (size, length, complexity, cohesion, coupling). It does not intend to be complete or fully objective; other frameworks could have been proposed and different choices could have been made. However, we believe that the formalisms and properties we introduce are convenient and intuitive. This framework contributes constructively to a firmer theoretical ground of software measurement.
Lionel C. Briand, Sandro Morasca, Victor R. Basili
IEEE Trans. Software Eng.2
1995 Generating Test Cases for Real-Time Systems from Logic Specifications
abstract
We address the problem of automated derivation of functional test cases for real-time systems, by introducing techniques for generating test cases from formal specifications written in TRIO, a language that extends classical temporal logic to deal explicitly with time measures. We describe an interactive tool that has been built to implement these techniques, based on interpretation algorithms of the TRIO language. Several heuristic criteria are suggested to reduce drastically the size of the test cases that are generated. Experience in the use of the tool on real-life cases is reported.
Dino Mandrioli, Sandro Morasca, Angelo Morzenti
ACM Trans. Comput. Syst.2
1994 Validating timing requirements for time basic net specifications
Carlo Ghezzi, Sandro Morasca, Mauro Pezzè
J. Syst. Softw.2
1993 Measuring and Assessing Maintainability at the End of High Level Design
abstract
Software architecture appears to be one of the main factors affecting software maintainability. Therefore, in order to be able to predict and assess maintainability early in the development process one needs to be able to measure the high-level design characteristics that affect the change process. To this end, a measurement approach based on precise assumptions derived from the change process is proposed. The change process is based on object-oriented design principles and is partially language independent. Metrics for cohesion, coupling, and visibility are defined in order to capture the difficulty of isolating, understanding, designing and validating changes.>
Lionel C. Briand, Sandro Morasca, Victor R. Basili
ICSM2
1991 Timed High-Level Nets
Sandro Morasca, Mauro Pezzè, Marco Trubian
Real Time Syst.1
1991 A Unified High-Level Petri Net Formalism for Time-Critical Systems
abstract
The authors introduce a high-level Petri net formalism-environment/relationship (ER) nets-which can be used to specify control, function, and timing issues. In particular, they discuss how time can be modeled via ER nets by providing a suitable axiomatization. They use ER nets to define a time notation that is shown to generalize most time Petri-net-based formalisms which appeared in the literature. They discuss how ER nets can be used in a specification support environment for a time-critical system and, in particular, the kind of analysis supported.>
Carlo Ghezzi, Dino Mandrioli, Sandro Morasca, Mauro Pezzè
IEEE Trans. Software Eng.3
1990 Extending software complexity metrics to concurrent programs
abstract
A metric for concurrent software is proposed based on an abstract model (Petri nets) as an extension of T.J. McCabe's (1976) cyclomatic number. As such, its focus is on the complexity of control flow. This metric is applied to the assessment of Ada programs, and an automatic method for its direct computation based on the inspection of Ada code is provided. It is pointed out, however, that wider experimentation is needed in order to better assess its effectiveness.>
Flavio De Paoli, Sandro Morasca
COMPSAC2
1989 Symbolic Execution of Concurrent Systems Using Petri Nets
Carlo Ghezzi, Dino Mandrioli, Sandro Morasca, Mauro Pezzè
Comput. Lang.3
1986 Software metrics: A critical evaluation and an application to Pascal
Roberto Lecciso, Stefano Mainetti, Sandro Morasca
Microprocessing and Microprogramming3