Gabriele Rotoloni

dblp:348/7813 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0003-2046-0090ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 5 since 2021
YearPublicationVenuePosition
2025 Critical Considerations on Effort-aware Software Defect Prediction Metrics
abstract
Background. Effort-aware metrics (EAMs) are widely used to evaluate the effectiveness of software defect prediction models, while accounting for the effort needed to analyze the software modules that are estimated defective. The usual underlying assumption is that this effort is proportional to the modules’ size measured in LOC. However, the research on module analysis (including code understanding, inspection, testing, etc.) suggests that module analysis effort may be better correlated to code attributes other than size. Aim. We investigate whether assuming that module analysis effort is proportional to other code metrics than LOC leads to different evaluations. Method. We show mathematically that the choice of the code measure used as the module effort driver crucially influences the resulting evaluations. To illustrate the practical consequences of this, we carried out a demonstrative empirical study, in which the same model was evaluated via EAMs, assuming that effort is proportional to either McCabe’s complexity or LOC. Results. The empirical study showed that EAMs depend on the underlying effort model, and can give quite different indications when effort is modeled differently. It is also apparent that the extent of these differences varies widely. Conclusions. Researchers and practitioners should be aware that the reliability of the indications provided by EAMs depend on the nature of the underlying effort model. The EAMs used until now appear to be actually size-aware, rather than effort-aware: when analysis effort does not depend on size, these EAMs can be misleading.
Luigi Lavazza, Gabriele Rotoloni, Sandro Morasca
EASE2
2025 A Replicated Study on Factors Affecting Software Understandability
Georgia M. Kapitsaki, Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
ENASE4
2025 Software Defect Prediction evaluation: New metrics based on the ROC curve
abstract
Context: ROC (Receiver Operating Characteristic) curves are widely used to represent how well fault-proneness models (e.g., probability models) classify software modules as faulty or non-faulty. AUC , the Area Under the ROC Curve, is usually used to quantify the overall discriminating power of a fault-proneness model. Alternative indicators proposed, e.g., RRA (Ratio of Relevant Areas), consider the area under a portion of a ROC curve. Each point of a ROC curve represents a binary classifier, obtained by setting a specified threshold on the fault-proneness model. Several performance metrics (Precision, Recall, the F-score, etc.) are used to assess a binary classifier. Objectives: We investigate the relationships linking “under the ROC curve area” indicators such as AUC and RRA to performance metrics. Methods: We study these relationships analytically. We introduce iso-PM ROC curves, whose points have the same value P M ¯ for a given performance metric PM. When evaluating a ROC curve, we identify the iso-PM curve with the same value of AUC or RRA . Its P M ¯ can be seen as a property of the ROC curve and fault-proneness model under evaluation. Results: There is an S-shaped relationship between P M ¯ and AUC for performance metrics that do not depend on the proportion ρ of faulty modules, i.e., dataset balancedness. ϕ (Matthews Correlation Coefficient) depends on ρ : with very imbalanced datasets, AUC appears over-optimistic and ϕ over-pessimistic. RRA defines the region of interest in terms of ρ , so all performance metrics depend on ρ . RRA is related to performance metrics via S-shaped curves. Conclusion: Our proposal helps gain a better quantitative understanding of the goodness of a ROC curve, especially in practically relevant regions of interest. Also, showing a ROC curve and iso-PM curves provides an intuitive perception of the goodness of a fault-proneness model.
Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
Inf. Softw. Technol.3
2023 On the Reliability of the Area Under the ROC Curve in Empirical Software Engineering
abstract
Binary classifiers are commonly used in software engineering research to estimate several software qualities, e.g., defectiveness or vulnerability. Thus, it is important to adequately evaluate how well binary classifiers perform, before they are used in practice. The Area Under the Curve (AUC) of Receiver Operating Characteristic curves has often been used to this end. However, AUC has been the target of some criticisms, so it is necessary to evaluate under what conditions and to what extent AUC can be a reliable performance metric.
Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
EASE3
2023 An Experience in the Evaluation of Fault Prediction
Luigi Lavazza, Sandro Morasca, Gabriele Rotoloni
PROFES (1)3