EDBT 2026 Demo / reviewers in the wild / expert
Davide Falessi
dblp:89/536
· DBLP profile ↗
47ranked-venue papers
23as first author
19since 2021 · last 2027
0000-0002-6340-0058ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 44 · 22 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Beyond literacy: Predicting interpretation correctness of visualizations with user traits, item difficulty, and Rasch scoresabstractData Visualization Literacy assessments are often administered via fixed sets of Data Visualization (DV) items, despite heterogeneity in how different people interpret the same DV. In this study, we predict Human Interpretation Correctness, i.e., whether a specific person will answer a DV item correctly, prior to their exposure to the target DV. We operationalize this as Predicting Human Interpretation Correctness (P-HIC), a binary classification task using 22 pre-exposure features spanning Human Profile, Human Performance, and Item difficulty (i.e., experts’ ratings and Rasch model). In an online survey, 1083 participants answered 32 DV items (eight DVs × four items), yielding 34,656 responses. Across 32 item-specific datasets, 10× 10-fold cross-validation shows that a Bagging ensemble of J48 decision trees, combined with feature selection, performs best, achieving a median AUC of 0.73 and a median kappa of 0.33. Feature analyses indicate that item difficulty estimated by the Rasch model dominates prediction, followed by experts’ ratings and prior correctness (increasing in importance across sessions), while profile features contribute little. These results suggest that pre-exposure misinterpretation risk can be estimated above chance and warrant future evaluation of assistive item-ranking strategies in simulated and real adaptive assessment settings. Davide Falessi, Silvia Golia, Angela Locoro, Manuel Mastrofini |
Inf. Process. Manag. | 1 |
| 2026 | Anticipating bugs: Ticket-level bug prediction and temporal proximity effectsabstractAbstract Software bugs significantly impact project time, budgets, and safety, motivating extensive research in bug prediction. The primary goal of bug prediction is to optimize testing efforts by focusing on software fragments, i.e., classes, methods, commits (i.e., Just-In-Time or JIT), or lines of code, most likely to be buggy. However, these predictions are made only after defects have already been introduced. Thus, the current bug prediction approaches support fixing rather than prevention. Motivated by the principle of "prevention is better than cure," the aim of this paper is to introduce and evaluate Ticket-Level Prediction (TLP), an approach to identify tickets that will introduce bugs once implemented. We analyze TLP at three temporal points, each point represents a ticket lifecycle stage: Open, In Progress, or Closed. We conjecture that: (1) TLP accuracy increases as tickets progress towards the closed stage due to improved feature reliability over time, and (2) the predictive power of features changes across these temporal points. Our TLP approach leverages 72 features belonging to seven different families: code, developer, external temperature, internal temperature, intrinsic, ticket to tickets, and JIT. Our TLP evaluation uses a sliding-window approach, balancing feature selection and three machine-learning bug prediction classifiers on about 10,000 tickets of two Apache open-source projects. Our results show that TLP accuracy increases with proximity, con- firming the expected trade-off between early prediction and accuracy. Regarding the prediction power of feature families, no single feature family dominates across stages; developer-centric signals are most informative early, whereas code and JIT metrics prevail near closure, and temperature-based features provide complementary value throughout. Our findings complement and extend the literature on bug prediction at the class, method, or commit level by showing that defect predic- tion can be effectively moved upstream, offering opportunities for risk-aware ticket triaging and developer assignment before any code is written. Daniele La Prova, Emanuele Gentili, Davide Falessi |
Empir. Softw. Eng. | 3 |
| 2025 | Fewer Draws, More Fun: Searching for Unbalanced Positions in Chess
Afro Ambanelli, Paolo Ciancarini, Angelo Di Iorio, Davide Falessi, Andrea Manzo, Massimo Venuto |
ICEC | 4 |
| 2025 | Ticket-Augmented Just-In-Time Defect Prediction
Emanuele Gentili, Daniele La Prova, Davide Falessi |
PROFES | 3 |
| 2025 | Practitioners' perceptions on requirements smellsabstractContext: Software specifications are usually written in natural language and may suffer from imprecision, ambiguity, and other quality issues, hereafter referred to as requirement smells. Requirement smells can hinder project development in many aspects, such as delays, reworks, and low customer satisfaction. From an industrial perspective, we want to focus our time and effort on identifying and preventing the requirement smells of high interest. We also want to identify the metrics to measure the effect of smells on a software project. Objective: We aim to characterise types of requirement smells in terms of frequency, severity, and effects. To the best of our knowledge, no previous study analysed how frequency, severity, or effects vary across types of smells. Methods: We interview ten experienced practitioners from different divisions of a large international company in the safety-critical domain called MBDA Italy Spa. Then we survey 58 people from the same company to support our findings and extend the analysis to metrics for measuring specific types of requirements smells effects. Results: Our results show that the smell types perceived as most severe are Ambiguity and Unverifiability, while the most frequent are Ambiguity and Incompleteness. We also provide six Findings about requirements smells, such as that the effects of smells are expected to differ across smell types and stages of the project. our study suggests that measuring the effects of requirement smells may necessitate type-specific metrics. Conclusion: Our results contribute to a greater understanding of the importance of addressing requirement smells and provide actionable insights for improving requirement quality in industrial settings. Our results pave the way for future empirical investigations, such as mining project repositories, to measure the specific effect type and size of specific requirements’ smells. Emanuele Gentili, Davide Falessi |
Inf. Softw. Technol. | 2 |
| 2024 | An Extensive Comparison of Static Application Security Testing ToolsabstractContext: Static Application Security Testing Tools (SASTTs) identify software vulnerabilities to support the security and reliability of software applications. Interestingly, several studies have suggested that alternative solutions may be more effective than SASTTs due to their tendency to generate false alarms, commonly referred to as low Precision. Aim: We aim to comprehensively evaluate SASTTs, setting a reliable benchmark for assessing and finding gaps in vulnerability identification mechanisms based on SASTTs or alternatives. Method: Our SASTTs evaluation is based on a controlled, though synthetic, Java codebase. It involves an assessment of 1.5 million test executions, and it features innovative methodological features such as effort-aware accuracy metrics and method-level analysis. Results: Our findings reveal that SASTTs detect a tiny range of vulnerabilities. In contrast to prevailing wisdom, SASTTs exhibit high Precision while falling short in Recall. Conclusions: Our findings suggest that enhancing Recall, alongside expanding the spectrum of detected vulnerability types, should be the primary focus for improving SASTTs or alternative approaches, such as machine learning-based vulnerability identification solutions. Matteo Esposito 0001, Valentina Falaschi, Davide Falessi |
EASE | 3 |
| 2024 | A Systematic Mapping Study on Impact Analysis
Emanuele Gentili, Jonida Çarka, Davide Falessi |
ICSOFT | 3 |
| 2024 | VALIDATE: A deep dive into vulnerability prediction datasetsabstractVulnerabilities are an essential issue today, as they cause economic damage to the industry and endanger our daily life by threatening critical national security infrastructures. Vulnerability prediction supports software engineers in preventing the use of vulnerabilities by malicious attackers, thus improving the security and reliability of software. Datasets are vital to vulnerability prediction studies, as machine learning models require a dataset. Dataset creation is time-consuming, error-prone, and difficult to validate. This study aims to characterise the datasets of prediction studies in terms of availability and features. Moreover, to support researchers in finding and sharing datasets, we provide the first VulnerAbiLty predIction DatAseT rEpository (VALIDATE). We perform a systematic literature review of the datasets of vulnerability prediction studies. Our results show that out of 50 primary studies, only 22 studies (i.e., 38%) provide a reachable dataset. Of these 22 studies, only one study provides a dataset in a stable repository. Our repository of 31 datasets, 22 reachable plus nine datasets provided by authors via email, supports researchers in finding datasets of interest, hence avoiding reinventing the wheel; this translates into less effort, more reliability, and more reproducibility in dataset creation and use. Matteo Esposito 0001, Davide Falessi |
Inf. Softw. Technol. | 2 |
| 2024 | A Comprehensive View on TD Prevention Practices and Reasons for Not Preventing ItabstractContext . Technical debt (TD) prevention allows software practitioners to apply practices to avoid potential TD items in their projects. Aims . To uncover and prioritize, from the point of view of software practitioners, the practices that could be used to avoid TD items, the relations between these practices and the causes of TD, and the practice avoidance reasons (PARs) that could explain the failure to prevent TD. Method . We analyze data collected from six replications of a global industrial family of surveys on TD, totaling 653 answers. We also conducted a follow up survey to understand the importance level of analyzed data. Results . Most practitioners indicated that TD could be prevented, revealing 89 prevention practices and 23 PARs for explaining the failure to prevent TD. The article identifies statistically significant relationships between preventive practices and certain causes of TD. Further, it prioritizes the list of practices, PARs, and relationships regarding their level of importance for TD prevention based on the opinion of software practitioners. Conclusion . This work organizes TD prevention practices and PARs in a conceptual map and the relationships between practices and causes of TD in a Sankey diagram to help the visualization of the body of knowledge reported in this study. Sávio Freire, Alexia Pacheco, Nicolli Rios, Boris Perez, Camilo Castellanos, Darío Correal, Robert Ramac, Vladimir Mandic, Nebojsa Tausan, Gustavo López 0001, Manoel G. Mendonça, Davide Falessi, Clemente Izurieta, Carolyn B. Seaman, Rodrigo O. Spínola |
ACM Trans. Softw. Eng. Methodol. | 12 |
| 2023 | Characterizing Requirements Smells
Emanuele Gentili, Davide Falessi |
PROFES (1) | 2 |
| 2023 | Uncovering the Hidden Risks: The Importance of Predicting Bugginess in Untouched MethodsabstractBugs in untouched code can be a ticking time bomb, but are they worth predicting? Our study dives deep into the importance of predicting bugginess in untouched methods and its impact on bug prediction accuracy. Our analysis of six open-source projects reveals the hidden risks of dormant bugs and the benefits of predicting in isolation untouched methods. Our findings significantly increase prediction accuracy, decrease train dataset sizes, and prove that untouched methods are an overlooked yet vital aspect of bug prediction. Matteo Esposito 0001, Davide Falessi |
SCAM | 2 |
| 2023 | Can We Trust the Default Vulnerabilities Severity?abstractAs software systems become increasingly complex and interconnected, the risk of security debt has risen significantly, increasing cyber-attacks and data breaches. Vulnerability prioritization is a critical activity in software engineering as it helps identify and address security vulnerabilities in software systems promptly and effectively. With the increasing complexity of software systems and the growing number of potential threats, it is essential to have a systematic approach to vulnerability prioritization to ensure that the most critical vulnerabilities are addressed first. The present study aims to investigate the agreement between the default and the National Vulnerability Database (NVD) severity levels. We analyzed 1626 vulnerabilities encompassing 12 unique types of vulnerabilities associated with 125 Common Platform Enumeration identifiers belonging to 105 Apache projects. Our results show a scarce correlation between the default and NVD severity levels. Thus, the default severity of vulnerabilities is not trustworthy. Moreover, we discovered that, surprisingly, the same type of vulnerability has several NVD severity; therefore, no default prioritization can be accurate based only on the type of vulnerability. Future studies are needed to accurately estimate the priority of vulnerabilities by considering several aspects of vulnerabilities rather than only the type. Matteo Esposito 0001, Sergio Moreschini, Valentina Lenarduzzi, David Hästbacka, Davide Falessi |
SCAM | 5 |
| 2023 | Enhancing the defectiveness prediction of methods and classes via JITabstractAbstract Context Defect prediction can help at prioritizing testing tasks by, for instance, ranking a list of items (methods and classes) according to their likelihood to be defective. While many studies investigated how to predict the defectiveness of commits, methods, or classes separately, no study investigated how these predictions differ or benefit each other. Specifically, at the end of a release, before the code is shipped to production, testing can be aided by ranking methods or classes, and we do not know which of the two approaches is more accurate. Moreover, every commit touches one or more methods in one or more classes; hence, the likelihood of a method and a class being defective can be associated with the likelihood of the touching commits being defective. Thus, it is reasonable to assume that the accuracy of methods-defectiveness-predictions (MDP) and the class-defectiveness-predictions (CDP) are increased by leveraging commits-defectiveness-predictions (aka JIT). Objective The contribution of this paper is fourfold: (i) We compare methods and classes in terms of defectiveness and (ii) of accuracy in defectiveness prediction, (iii) we propose and evaluate a first and simple approach that leverages JIT to increase MDP accuracy and (iv) CDP accuracy. Method We analyse accuracy using two types of metrics (threshold-independent and effort-aware). We also use feature selection metrics, nine machine learning defect prediction classifiers, more than 2.000 defects related to 38 releases of nine open source projects from the Apache ecosystem. Our results are based on a ground truth with a total of 285,139 data points and 46 features among commits, methods and classes. Results Our results show that leveraging JIT by using a simple median approach increases the accuracy of MDP by an average of 17% AUC and 46% PofB10 while it increases the accuracy of CDP by an average of 31% AUC and 38% PofB20. Conclusions From a practitioner’s perspective, it is better to predict and rank defective methods than defective classes. From a researcher’s perspective, there is a high potential for leveraging statement-defectiveness-prediction (SDP) to aid MDP and CDP. Davide Falessi, Simone Mesiano Laureani, Jonida Çarka, Matteo Esposito 0001, Daniel Alencar da Costa |
Empir. Softw. Eng. | 1 |
| 2023 | Software practitioners' point of view on technical debt payment
Sávio Freire, Nicolli Rios, Boris Perez, Camilo Castellanos, Darío Correal, Robert Ramac, Vladimir Mandic, Nebojsa Tausan, Gustavo López 0001, Alexia Pacheco, Manoel G. Mendonça, Davide Falessi, Clemente Izurieta, Carolyn B. Seaman, Rodrigo O. Spínola |
J. Syst. Softw. | 12 |
| 2022 | On effort-aware metrics for defect predictionabstractAbstract Context Advances in defect prediction models, aka classifiers, have been validated via accuracy metrics. Effort-aware metrics (EAMs) relate to benefits provided by a classifier in accurately ranking defective entities such as classes or methods. PofB is an EAM that relates to a user that follows a ranking of the probability that an entity is defective, provided by the classifier. Despite the importance of EAMs, there is no study investigating EAMs trends and validity. Aim The aim of this paper is twofold: 1) we reveal issues in EAMs usage, and 2) we propose and evaluate a normalization of PofBs (aka NPofBs), which is based on ranking defective entities by predicted defect density. Method We perform a systematic mapping study featuring 152 primary studies in major journals and an empirical study featuring 10 EAMs, 10 classifiers, two industrial, and 12 open-source projects. Results Our systematic mapping study reveals that most studies using EAMs use only a single EAM (e.g., PofB20) and that some studies mismatched EAMs names. The main result of our empirical study is that NPofBs are statistically and by orders of magnitude higher than PofBs. Conclusions In conclusion, the proposed normalization of PofBs: (i) increases the realism of results as it relates to a better use of classifiers, and (ii) promotes the practical adoption of prediction models in industry as it shows higher benefits. Finally, we provide a tool to compute EAMs to support researchers in avoiding past issues in using EAMs. Jonida Çarka, Matteo Esposito 0001, Davide Falessi |
Empir. Softw. Eng. | 3 |
| 2022 | The Impact of Dormant Defects on Defect Prediction: A Study of 19 Apache ProjectsabstractDefect prediction models can be beneficial to prioritize testing, analysis, or code review activities, and has been the subject of a substantial effort in academia, and some applications in industrial contexts. A necessary precondition when creating a defect prediction model is the availability of defect data from the history of projects. If this data is noisy, the resulting defect prediction model could result to be unreliable. One of the causes of noise for defect datasets is the presence of “dormant defects,” i.e., of defects discovered several releases after their introduction. This can cause a class to be labeled as defect-free while it is not, and is, therefore “snoring.” In this article, we investigate the impact of snoring on classifiers' accuracy and the effectiveness of a possible countermeasure, i.e., dropping too recent data from a training set. We analyze the accuracy of 15 machine learning defect prediction classifiers, on data from more than 4,000 defects and 600 releases of 19 open source projects from the Apache ecosystem. Our results show that on average across projects (i) the presence of dormant defects decreases the recall of defect prediction classifiers, and (ii) removing from the training set the classes that in the last release are labeled as not defective significantly improves the accuracy of the classifiers. In summary, this article provides insights on how to create defects datasets by mitigating the negative effect of dormant defects on defect prediction. Davide Falessi, Aalok Ahluwalia, Massimiliano Di Penta |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2021 | Worst Smells and Their Worst ReasonsabstractCode bad smells are symptoms of poor design and implementation. There are several well-known smell types, such as large classes (aka God classes), code clones, etc. and they have been shown to lead to technical debt and hence to decrease code maintainability. Quality gates are a recent technology that prevents the automatic acceptance of push requests of code commits that have been identified as containing certain smells. However, it is a challenging activity to decide which smells should be included in the quality gate, as developers may choose to optimize short term benefits like time to market over long term benefits like maintainability. But some smells appear to provide no benefit to developers whatsoever and hence such smells should always be avoided. The aims of this paper are: 1) to identify "worst smells", i.e., bad smells that never have a good reason to exist, 2) to determine the frequency, change-proneness, and severity associated with worst smells, and 3) to identify the "worst reasons", i.e., the reasons for introducing these worst smells in the first place. To achieve these aims we ran a survey with 71 developers. We learned that 80 out of 314 catalogued code smells are "worst"; that is, developers agreed that these 80 smells should never exist in any code base. We then checked the frequency and change-proneness of these worst smells on 27 large Apache open- source projects. Our results show insignificant differences, in both frequency and change proneness, between worst and non-worst smells. That is to say, these smells are just as damaging as other smells, but there is never any justifiable reason to introduce them. Finally, in follow-up phone interviews with five developers we confirmed that these smells are indeed worst, and the interviewees proposed seven reasons for why they may be introduced in the first place. By explicitly identifying these seven reasons, project stakeholders can, through quality gates or reviews, ensure that such smells are never accepted in a code base, thus improving quality without compromising other goals such as agility or time to market. Davide Falessi, Rick Kazman |
TechDebt@ICSE | 1 |
| 2021 | Leveraging Intermediate Artifacts to Improve Automated Trace Link RetrievalabstractSoftware traceability establishes a network of connections between diverse artifacts such as requirements, design, and code. However, given the cost and effort of creating and maintaining trace links manually, researchers have proposed automated approaches using information retrieval techniques. Current approaches focus almost entirely upon generating links between pairs of artifacts and have not leveraged the broader network of interconnected artifacts. In this paper we investigate the use of intermediate artifacts to enhance the accuracy of the generated trace links - focusing on paths consisting of source, target, and intermediate artifacts. We propose and evaluate combinations of techniques for computing semantic similarity, scaling scores across multiple paths, and aggregating results from multiple paths. We report results from five projects, including one large industrial project. We find that leveraging intermediate artifacts improves the accuracy of end-to-end trace retrieval across all datasets and accuracy metrics. After further analysis, we discover that leveraging intermediate artifacts is only helpful when a project's artifacts share a common vocabulary, which tends to occur in refinement and decomposition hierarchies of artifacts. Given our hybrid approach that integrates both direct and transitive links, we observed little to no loss of accuracy when intermediate artifacts lacked a shared vocabulary with source or target artifacts. Alberto D. Rodriguez, Jane Cleland-Huang, Davide Falessi |
ICSME | 3 |
| 2021 | Leveraging the Defects Life Cycle to Label Affected Versions and Defective ClassesabstractTwo recent studies explicitly recommend labeling defective classes in releases using the affected versions (AV) available in issue trackers (e.g., Jira). This practice is coined as the realistic approach . However, no study has investigated whether it is feasible to rely on AVs. For example, how available and consistent is the AV information on existing issue trackers? Additionally, no study has attempted to retrieve AVs when they are unavailable. The aim of our study is threefold: (1) to measure the proportion of defects for which the realistic method is usable, (2) to propose a method for retrieving the AVs of a defect, thus making the realistic approach usable when AVs are unavailable, (3) to compare the accuracy of the proposed method versus three SZZ implementations. The assumption of our proposed method is that defects have a stable life cycle in terms of the proportion of the number of versions affected by the defects before discovering and fixing these defects. Results related to 212 open-source projects from the Apache ecosystem, featuring a total of about 125,000 defects, reveal that the realistic method cannot be used in the majority (51%) of defects. Therefore, it is important to develop automated methods to retrieve AVs. Results related to 76 open-source projects from the Apache ecosystem, featuring a total of about 6,250,000 classes, affected by 60,000 defects, and spread over 4,000 versions and 760,000 commits, reveal that the proportion of the number of versions between defect discovery and fix is pretty stable (standard deviation <2)—across the defects of the same project. Moreover, the proposed method resulted significantly more accurate than all three SZZ implementations in (i) retrieving AVs, (ii) labeling classes as defective, and (iii) in developing defects repositories to perform feature selection. Thus, when the realistic method is unusable, the proposed method is a valid automated alternative to SZZ for retrieving the origin of a defect. Finally, given the low accuracy of SZZ, researchers should consider re-executing the studies that have used SZZ as an oracle and, in general, should prefer selecting projects with a high proportion of available and consistent AVs. Bailey Vandehei, Daniel Alencar da Costa, Davide Falessi |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2020 | On the need of preserving order of data when validating within-project defect classifiersabstractAbstract We are in the shoes of a practitioner who uses previous project releases’ data to predict which classes of the current release are defect-prone. In this scenario, the practitioner would like to use the most accurate classifier among the many available ones. A validation technique, hereinafter “technique”, defines how to measure the prediction accuracy of a classifier. Several previous research efforts analyzed several techniques. However, no previous study compared validation techniques in the within-project across-release class-level context or considered techniques that preserve the order of data. In this paper, we investigate which technique recommends the most accurate classifier. We use the last release of a project as the ground truth to evaluate the classifier’s accuracy and hence the ability of a technique to recommend an accurate classifier. We consider nine classifiers, two industry and 13 open projects, and three validation techniques: namely 10-fold cross-validation (i.e., the most used technique), bootstrap (i.e., the recommended technique), and walk-forward (i.e., a technique preserving the order of data). Our results show that: 1) classifiers differ in accuracy in all datasets regardless of their entity per value, 2) walk-forward outperforms both 10-fold cross-validation and bootstrap statistically in all three accuracy metrics: AUC of the selected classifier, bias and absolute bias, 3) surprisingly, all techniques resulted to be more prone to overestimate than to underestimate the performances of classifiers, and 3) the defect rate resulted in changing between the second and first half in both industry projects and 83% of open-source datasets. This study recommends the use of techniques that preserve the order of data such as walk-forward over 10-fold cross-validation and bootstrap in the within-project across-release class-level context given the above empirical results and that walk-forward is by nature more simple, inexpensive, and stable than the other two techniques. Davide Falessi, Jacky Huang, Likhita Narayana, Jennifer Fong Thai, Burak Turhan |
Empir. Softw. Eng. | 1 |
| 2020 | Correction to: On the need of preserving order of data when validating within-project defect classifiersabstractTo fulfill the contractual requirement of the Compact agreement, the following funding note has to be added and placed in the Funding section of the original article: Open access funding provided by Università degli Studi di Roma Tor Vergata within the CRUI-CARE Agreement. Davide Falessi, Jacky Huang, Likhita Narayana, Jennifer Fong Thai, Burak Turhan |
Empir. Softw. Eng. | 1 |
| 2020 | Leveraging Historical Associations between Requirements and Source Code to Identify Impacted ClassesabstractAs new requirements are introduced and implemented in a software system, developers must identify the set of source code classes which need to be changed. Therefore, past effort has focused on predicting the set of classes impacted by a requirement. In this paper, we introduce and evaluate a new type of information based on the intuition that the set of requirements which are associated with historical changes to a specific class are likely to exhibit semantic similarity to new requirements which impact that class. This new Requirements to Requirements Set (R2RS) family of metrics captures the semantic similarity between a new requirement and the set of existing requirements previously associated with a class. The aim of this paper is to present and evaluate the usefulness of R2RS metrics in predicting the set of classes impacted by a requirement. We consider 18 different R2RS metrics by combining six natural language processing techniques to measure the semantic similarity among texts (e.g., VSM) and three distribution scores to compute overall similarity (e.g., average among similarity scores). We evaluate if R2RS is useful for predicting impacted classes in combination and against four other families of metrics that are based upon temporal locality of changes, direct similarity to code, complexity metrics, and code smells. Our evaluation features five classifiers and 78 releases belonging to four large open-source projects, which result in over 700,000 candidate impacted classes. Experimental results show that leveraging R2RS information increases the accuracy of predicting impacted classes practically by an average of more than 60 percent across the various classifiers and projects. Davide Falessi, Justin Roll, Jin L. C. Guo, Jane Cleland-Huang |
IEEE Trans. Software Eng. | 1 |
| 2019 | Snoring: a noise in defect prediction datasetsabstractIn order to develop and train defect prediction models, researchers rely on datasets in which a defect is often attributed to a release where the defect itself is discovered. However, in many circumstances, it can happen that a defect is only discovered several releases after its introduction. This might introduce a bias in the dataset, i.e., treating the intermediate releases as defect-free and the latter as defect-prone. We call this phenomenon as "sleeping defects". We call "snoring" the phenomenon where classes are affected by sleeping defects only, that would be treated as defect-free until the defect is discovered. In this paper we analyze, on data from 282 releases of six open source projects from the Apache ecosystem, the magnitude of the sleeping defects and of the snoring classes. Our results indicate that 1) on all projects, most of the defects in a project slept for more than 20% of the existing releases, and 2) in the majority of the projects the missing rate is more than 25% even if we remove 50% of releases. Aalok Ahluwalia, Davide Falessi, Massimiliano Di Penta |
MSR | 2 |
| 2018 | A Reflection on Diversity and Inclusivity Efforts in a Software Engineering ProgramabstractThis Innovative Practice Full Paper is an experience report that presents a collection of initiatives, circumstances, and teaching practices that coincide with improvements in gender diversity in the undergraduate software engineering program at California Polytechnic State University at San Luis Obispo. The percent of females in the software engineering capstone increased from an average of 3.81% in the first five years of the program (2003-2007) to 18.82% in the most recent four years (2014-2017). Multiple initiatives were instituted beginning in 2009 to improve a gender imbalance, addressing recruitment and retention of women. One key initiative was the creation of a new introductory course with multiple themes (e.g. art, mobile, music, robotics) from which incoming students could choose. These courses were designed to include significant collaboration, rapid application development in interesting domains, and strategic selection of tools and languages that reduced the advantages of previous student programming experience. Additional initiatives included club activities, sending large numbers of female students to the Grace Hopper Celebration of Women in Computing Conference, K-12 outreach, and a vibrant and mature SE capstone experience. Also during this timeframe, course scheduling changes were imposed which naturally created informal cohorts of software engineering students earlier than in previous years. Student self-evaluations collected in the SE capstone were analyzed, comparing male and female responses, as well as teams with different gender mixes. This analysis indicates no significant difference between male and female enjoyment of the capstone projects overall, and no significant difference between team enjoyment regardless of the percentage of females on the team. David S. Janzen, Sara Bahrami, Bruno da Silva 0002, Davide Falessi |
FIE | 4 |
| 2018 | Issue Tracking Systems: What Developers Want and Use
Davide Falessi, Freddy Hernandez, Foaad Khosmood |
ICSOFT | 1 |
| 2018 | Estimating the number of remaining links in traceability recovery (journal-first abstract)abstractAlthough very important in software engineering, establishing traceability links between software artifacts is extremely tedious, error-prone, and it requires significant effort. Even when approaches for automated traceability recovery exist, these provide the requirements analyst with a, usually very long, ranked list of candidate links that needs to be manually inspected. In this paper we introduce an approach called Estimation of the Number of Remaining Links (ENRL) which aims at estimating, via Machine Learning (ML) classifiers, the number of remaining positive links in a ranked list of candidate traceability links produced by a Natural Language Processing techniques-based recovery approach. We have evaluated the accuracy of the ENRL approach by considering several ML classifiers and NLP techniques on three datasets from industry and academia, and concerning traceability links among different kinds of software artifacts including requirements, use cases, design documents, source code, and test cases. Results from our study indicate that: (i) specific estimation models are able to provide accurate estimates of the number of remaining positive links; (ii) the estimation accuracy depends on the choice of the NLP technique, and (iii) univariate estimation models outperform multivariate ones. Davide Falessi, Massimiliano Di Penta, Gerardo Canfora, Giovanni Cantone |
ASE | 1 |
| 2018 | Empirical software engineering experts on the use of students and professionals in experiments
Davide Falessi, Natalia Juristo Juzgado, Claes Wohlin, Burak Turhan, Jürgen Münch, Andreas Jedlitschka, Markku Oivo |
Empir. Softw. Eng. | 1 |
| 2018 | Four commentaries on the use of students and professionals in empirical software engineering experiments
Robert Feldt, Thomas Zimmermann 0001, Gunnar R. Bergersen, Davide Falessi, Andreas Jedlitschka, Natalia Juristo Juzgado, Jürgen Münch, Markku Oivo, Per Runeson, Martin J. Shepperd, Dag I. K. Sjøberg, Burak Turhan |
Empir. Softw. Eng. | 4 |
| 2017 | What if I Had No Smells?abstractWhat would have happened if I did not have any code smell? This is an interesting question that no previous study, to the best of our knowledge, has tried to answer. In this paper, we present a method for implementing a what-if scenario analysis estimating the number of defective files in the absence of smells. Our industrial case study shows that 20% of the total defective files were likely avoidable by avoiding smells. Such estimation needs to be used with the due care though as it is based on a hypothetical history (i.e., zero number of smells and same process and product change characteristics). Specifically, the number of defective files could even increase for some types of smells. In addition, we note that in some circumstances, accepting code with smells might still be a good option for a company. Davide Falessi, Barbara Russo, Kathleen Mullen |
ESEM | 1 |
| 2017 | STRESS: A Semi-Automated, Fully Replicable Approach for Project SelectionabstractThe mining of software repositories has provided significant advances in a multitude of software engineering fields, including defect prediction. Several studies show that the performance of a software engineering technology (e.g., prediction model) differs across different project repositories. Thus, it is important that the project selection is replicable. The aim of this paper is to present STRESS, a semi-automated and fully replicable approach that allows researchers to select projects by configuring the desired level of diversity, fit, and quality. STRESS records the rationale behind the researcher decisions and allows different users to re-run or modify such decisions. STRESS is open-source and it can be used used locally or even online (www.falessi.com/STRESS/). We perform a systematic mapping study that considers studies that analyzed projects managed with JIRA and Git to asses the project selection replicability of past studies. We validate the feasible application of STRESS in realistic research scenarios by applying STRESS to select projects among the 211 Apache Software Foundation projects. Our systematic mapping study results show that none of the 68 analyzed studies is completely replicable. Regarding STRESS, it successfully supported the project selection among all 211 ASF projects. It also supported the measurement of 100 projects characteristics, including the 32 criteria of the studies analyzed in our mapping study. The mapping study and STRESS are, to our best knowledge, the first attempt to investigate and support the replicability of project selection. We plan to extend them to other technologies such as GitHub. Davide Falessi, Wyatt Smith, Alexander Serebrenik |
ESEM | 1 |
| 2017 | Estimating the number of remaining links in traceability recovery
Davide Falessi, Massimiliano Di Penta, Gerardo Canfora, Giovanni Cantone |
Empir. Softw. Eng. | 1 |
| 2016 | Introduction to the special issue on technical debt in software systemsabstractFor technical reasons the number of authors shown on this cover page is limited to 10 maximum. Davide Falessi, Philippe Kruchten, Paris Avgeriou |
J. Syst. Softw. | 1 |
| 2015 | A Mapping Study of Software Causal Factors for Improving MaintenanceabstractContext: Software maintenance is important to keep existing software systems functional for organizations or users that depend on that software. Goal: We aim to identify the factors, i.e., software characteristics such as code complexity, leading to maintenance problems. Method: We present a Mapping Study (MS) on controlled experiments that investigated software characteristics related to defects during maintenance. Results: The search strategy identified 78 papers, of which 9 have been included in our study, dated from 1985 to 2013, after applying our inclusion and exclusion criteria. We extracted data from these papers to identify the research methods, and the independent, dependent, blocked, and measured variables. Conclusions: Our MS results point to a weak evidence on software factors causing defects during maintenance. Stronger evidence can be developed via more controlled experiments that address multiple independent variables and hold the software objects constant. Carson Carroll, Davide Falessi, Vanessa Forney, Alexa Frances, Clemente Izurieta, Carolyn B. Seaman |
ESEM | 2 |
| 2015 | Evidence management for compliance of critical systems with safety standards: A survey on the state of practice
Sunil Nair, Jose Luis de la Vara, Mehrdad Sabetzadeh, Davide Falessi |
Inf. Softw. Technol. | 4 |
| 2014 | Traceability and SysML design slices to support safety inspections: A controlled experimentabstractCertifying safety-critical software and ensuring its safety requires checking the conformance between safety requirements and design. Increasingly, the development of safety-critical software relies on modeling, and the System Modeling Language (SysML) is now commonly used in many industry sectors. Inspecting safety conformance by comparing design models against safety requirements requires safety inspectors to browse through large models and is consequently time consuming and error-prone. To address this, we have devised a mechanism to establish traceability between (functional) safety requirements and SysML design models to extract design slices (model fragments) that filter out irrelevant details but keep enough context information for the slices to be easy to inspect and understand. In this article, we report on a controlled experiment assessing the impact of the traceability and slicing mechanism on inspectors' conformance decisions and effort. Results show a significant decrease in effort and an increase in decisions' correctness and level of certainty. Lionel C. Briand, Davide Falessi, Shiva Nejati 0001, Mehrdad Sabetzadeh, Tao Yue 0002 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2013 | Message from the MTD 2013 Workshop ChairsabstractThe goal of this fifth workshop on managing technical debt is to engage leading empiricists and practitioners in exploring practical problems to provide opportunities for research that can provide evidence for or against the emerging definition of technical debt and the efficacy of emerging practice. Philippe Kruchten, Robert L. Nord, Ipek Ozkaya, Davide Falessi |
ESEM | 4 |
| 2013 | Appendix: ICSR 2013 Workshop Summaries
Davide Falessi |
ICSR | 1 |
| 2013 | The value of design rationale informationabstractA complete and detailed (full) Design Rationale Documentation (DRD) could support many software development activities, such as an impact analysis or a major redesign. However, this is typically too onerous for systematic industrial use as it is not cost effective to write, maintain, or read. The key idea investigated in this article is that DRD should be developed only to the extent required to support activities particularly difficult to execute or in need of significant improvement in a particular context. The aim of this article is to empirically investigate the customization of the DRD by documenting only the information items that will probably be required for executing an activity. This customization strategy relies on the hypothesis that the value of a specific DRD information item depends on its category (e.g., assumptions, related requirements, etc.) and on the activity it is meant to support. We investigate this hypothesis through two controlled experiments involving a total of 75 master students as experimental subjects. Results show that the value of a DRD information item significantly depends on its category and, within a given category, on the activity it supports. Furthermore, on average among activities, documenting only the information items that have been required at least half of the time (i.e., the information that will probably be required in the future) leads to a customized DRD containing about half the information items of a full documentation. We expect that such a significant reduction in DRD information should mitigate the effects of some inhibitors that currently prevent practitioners from documenting design decision rationale. Davide Falessi, Lionel C. Briand, Giovanni Cantone, Rafael Capilla, Philippe Kruchten |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2013 | Empirical Principles and an Industrial Case Study in Retrieving Equivalent Requirements via Natural Language Processing TechniquesabstractThough very important in software engineering, linking artifacts of the same type (clone detection) or different types (traceability recovery) is extremely tedious, error-prone, and effort-intensive. Past research focused on supporting analysts with techniques based on Natural Language Processing (NLP) to identify candidate links. Because many NLP techniques exist and their performance varies according to context, it is crucial to define and use reliable evaluation procedures. The aim of this paper is to propose a set of seven principles for evaluating the performance of NLP techniques in identifying equivalent requirements. In this paper, we conjecture, and verify, that NLP techniques perform on a given dataset according to both ability and the odds of identifying equivalent requirements correctly. For instance, when the odds of identifying equivalent requirements are very high, then it is reasonable to expect that NLP techniques will result in good performance. Our key idea is to measure this random factor of the specific dataset(s) in use and then adjust the observed performance accordingly. To support the application of the principles we report their practical application to a case study that evaluates the performance of a large number of NLP techniques for identifying equivalent requirements in the context of an Italian company in the defense and aerospace domain. The current application context is the evaluation of NLP techniques to identify equivalent requirements. However, most of the proposed principles seem applicable to evaluating any estimation technique aimed at supporting a binary decision (e.g., equivalent/nonequivalent), with the estimate in the range [0,1] (e.g., the similarity provided by the NLP), when the dataset(s) is used as a benchmark (i.e., testbed), independently of the type of estimator (i.e., requirements text) and of the estimation method (e.g., NLP). Davide Falessi, Giovanni Cantone, Gerardo Canfora |
IEEE Trans. Software Eng. | 1 |
| 2012 | Research-Based Innovation: A Tale of Three Projects in Model-Driven Engineering
Lionel C. Briand, Davide Falessi, Shiva Nejati 0001, Mehrdad Sabetzadeh, Tao Yue 0002 |
MoDELS | 2 |
| 2012 | A SysML-based approach to traceability management and design slicing in support of safety certification: Framework, tool support, and case studies
Shiva Nejati 0001, Mehrdad Sabetzadeh, Davide Falessi, Lionel C. Briand, Thierry Coq |
Inf. Softw. Technol. | 3 |
| 2011 | SafeSlice: a model slicing and design safety inspection tool for SysMLabstractSoftware safety certification involves checking that the software design meets the (software) safety requirements. In practice, inspections are one of the primary vehicles for ensuring that safety requirements are satisfied by the design. Unless the safety-related aspects of the design are clearly delineated, the inspections conducted by safety assessors would have to consider the entire design, although only small fragments of the design may be related to safety. In a model-driven development context, this means that the assessors have to browse through large models, understand them, and identify the safety-related fragments. This is time-consuming and error-prone, specially noting that the assessors are often third-party regulatory bodies who were not involved in the design. To address this problem, we describe in this paper a prototype tool called, SafeSlice, that enables one to automatically extract the safety-related slices (fragments) of design models. The main enabler for our slicing technique is the traceability between the safety requirements and the design, established by following a structured design methodology that we propose. Our work is grounded on SysML, which is being increasingly used for expressing the design of safety-critical systems. We have validated our work through two case studies and a control experiment which we briefly outline in the paper. Davide Falessi, Shiva Nejati 0001, Mehrdad Sabetzadeh, Lionel C. Briand, Antonio Messina |
SIGSOFT FSE | 1 |
| 2010 | A comprehensive characterization of NLP techniques for identifying equivalent requirementsabstractThough very important in software engineering, linking artifacts of the same type (clone detection) or of different types (traceability recovery) is extremely tedious, error-prone and requires significant effort. Past research focused on supporting analysts with mechanisms based on Natural Language Processing (NLP) to identify candidate links. Because a plethora of NLP techniques exists, and their performances vary among contexts, it is important to characterize them according to the provided level of support. The aim of this paper is to characterize a comprehensive set of NLP techniques according to the provided level of support to human analysts in detecting equivalent requirements. The characterization consists on a case study, featuring real requirements, in the context of an Italian company in the defense and aerospace domain. The major result from the case study is that simple NLP are more precise than complex ones. Davide Falessi, Giovanni Cantone, Gerardo Canfora |
ESEM | 1 |
| 2010 | Applying empirical software engineering to software architecture: challenges and lessons learned
Davide Falessi, Muhammad Ali Babar 0001, Giovanni Cantone, Philippe Kruchten |
Empir. Softw. Eng. | 1 |
| 2008 | Implementing Product Line Engineering in Industry: Feedback from the Field to Research
Davide Falessi, Dirk Muthig |
PROFES | 1 |
| 2007 | Issues in Applying Empirical Software Engineering to Software Architecture
Davide Falessi, Philippe Kruchten, Giovanni Cantone |
ECSA | 1 |
| 2007 | Do Architecture Design Methods Meet Architects' Needs?abstractSeveral Software Architecture Design Methods (SADM) have been published, reviewed, and compared. But these surveys and comparisons are mostly centered on intrinsic elements of the design method, and they do not compare them from the perspective of the actual needs of software architects. We would like to analyze the completeness of SADM from an architect's point of view. To do so, we define nine categories of software architects' needs, propose an ordinal scale for evaluating the degree to which a given SADM meets the needs, and then apply this to a small set of SADMs. The contribution of the paper is twofold: (i) to provide a different and useful frame of reference for architects to select SADM, and (ii) to suggest SADM areas of improvements. We found two answers to our question: "do architectural design methods meet the needs of the architect?" Yes, all architect's needs are met by one or another SADM, but No, no architectural design method meets simultaneously all the needs of an architect. This approach may lead to improvements of existing SADMs. Davide Falessi, Giovanni Cantone, Philippe Kruchten |
WICSA | 1 |