EDBT 2026 Demo / reviewers in the wild / expert
Carmine Gravino
dblp:62/1673
· DBLP profile ↗
82ranked-venue papers
5as first author
23since 2021 · last 2027
0000-0002-4394-9035ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 72 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6Theory of computation · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | From pixels to attribution reports: Diabetic retinopathy grading with CNN-transformer ensembles and vision-language modelsabstractDiabetic retinopathy (DR) grading from fundus images remains challenging because severity labels are ordinal, class distributions are imbalanced, and neighboring stages can be visually ambiguous. This study presents a unified pipeline for five-class DR grading, including controlled CNN/Transformer backbone benchmarking, ensemble evaluation, computational-cost analysis, external validation, and proof-of-concept attribution reporting. APTOS 2019 was used to evaluate six backbones under stratified five-fold cross-validation with fixed fold assignments, identical preprocessing, imbalance-aware training, and model selection based on quadratic weighted kappa (QWK). Hard voting, weighted soft voting, stacking, and hybrid class-level fusion were evaluated to combine complementary models, while trainable parameters, model size, inference latency, and throughput were reported to quantify deployment cost. Model generalization was assessed by directly evaluating APTOS-trained models on Messidor-2 without retraining, fine-tuning, or recalibration. Safety-filtered summaries generated with a vision-language model were restricted to non-diagnostic model-attribution descriptions, and fundus-masked Grad-CAM++ maps were paired with these summaries for attribution reporting. Among individual backbones, ResNet-50 and ConvNeXt-Tiny achieved the strongest performance, while weighted soft voting provided the most stable ensemble performance. External validation confirmed the difficulty of cross-dataset five-class DR staging, with reduced performance compared with internal cross-validation. The generated summaries adhered to predefined safety constraints but were not used for lesion localization, diagnosis, or clinical validation. Overall, the study demonstrates that controlled benchmarking and weighted ensembling can provide stable ordinal DR grading performance, while computational overhead, domain shift, uncertainty calibration, and expert clinical validation remain important considerations for practical deployment. Pir Bakhsh Khokhar, Carmine Gravino, Fabio Palomba, Sule Yildirim Yayilgan, Sarang Shaikh |
Expert Syst. Appl. | 2 |
| 2026 | RobustDRNet: A clinically-aligned hybrid ensemble model with multi-method explainability for lesion-aware diabetic retinopathy gradingabstractDiabetic retinopathy (DR) screening requires artificial intelligence (AI) models that are not only highly accurate in grading five clinical stages but are also capable of generating quantitatively evaluated lesion-aware explanations to earn the trust of clinicians. We propose RobustDRNet , a hybrid ensemble model that combines local convolutional features from Residual Network-34 (ResNet-34) and ConvNeXt-Tiny with global transformer embeddings from Vision Transformer Base/16 (ViT-B16) via two-stage feature fusion and a disentangled multilayer perceptron (MLP), followed by a logistic regression stacking meta-learner for prediction aggregation. To address severe class imbalance, our training pipeline employs stratified sampling, contrast-limited adaptive histogram equalization (CLAHE) for contrast enhancement, strong data augmentation, and class-weighted focal loss. Evaluated on the Asia Pacific Tele-Ophthalmology Society (APTOS) 2019 dataset, RobustDRNet achieved 88.4% validation accuracy, a 0.967 macro-averaged area under the receiver operating characteristic curve (macro-AUC), and Cohen’s kappa of 0.823, outperforming individual backbones and simple voting ensembles. In addition to classification performance, we integrated six complementary explainable AI (XAI) techniques: Gradient-weighted Class Activation Mapping++ (Grad-CAM++), Integrated Gradients, attention rollout, SHapley Additive exPlanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), and Testing with Concept Activation Vectors (TCAV). Each technique was quantitatively benchmarked against expert-annotated lesion maps from the Indian Diabetic Retinopathy Image Dataset (IDRiD). Saliency maps achieved mean Intersection over Union (IoU) scores of 0.06 for Grad-CAM++ and approximately 0.10 for Integrated Gradients; SHapley Additive exPlanations (SHAP) perturbations showed a deletion drop of 0.25 and an insertion gain of 0.22; and TCAV achieved complete classifier-level TCAV alignment (score = 1.0) with clinically coherent, grade-wise importance trajectories. By combining competitive grading performance with multi-perspective and quantitatively evaluated interpretability, RobustDRNet provides a promising DR screening framework whose decisions are supported by lesion-aware explanatory evidence. Pir Bakhsh Khokhar, Viviana Pentangelo, Carmine Gravino, Fabio Palomba |
Expert Syst. Appl. | 3 |
| 2026 | Automatic identification of privacy and security requirements: a systematic literature reviewabstractAbstract The utmost importance of privacy and security requirements in software development calls for adopting methods that enable the identification and proactive mitigation of these issues during the system development. Our survey of 45 primary studies provides an overview of the methods, document types, and datasets employed in tackling this challenge, along with an analysis of approaches demonstrating superior performance based on document types and specific identification problems. Analysis reveals a wide adoption of ML-based systems on diverse datasets, showcasing the effectiveness of leveraging various sources of information to identify privacy and security requirements in software development. Francesco Casillo, Vincenzo Deufemia, Carmine Gravino |
Requir. Eng. | 3 |
| 2025 | Towards Generating the Rationale for Code ChangesabstractCommit messages are essential to understand changes in software projects, providing a way for developers to communicate code evolution. Generating effective commit messages that explain the rationale behind changes is a challenging and timeconsuming task. While previous research has shown success in automating straightforward commit messages (e.g., “add README”), our study explores a more complex task: generating rationale explanations for code changes. We developed a method to identify rationale sentences in commit messages and compiled a dataset of 45,945 commits with their corresponding rationales. A pre-trained model was trained on this dataset to generate rationale explanations. While the approach we engineered for the extraction of rationale from commit messages exhibited a 75 % precision, the model trained to generate the rationale only worked in a minority of cases. Our findings highlight the difficulty of the tackled task and the need for additional research in the area. We release our dataset and code to foster the investigation of this problem. Francesco Casillo, Antonio Mastropaolo, Gabriele Bavota, Vincenzo Deufemia, Carmine Gravino |
ICPC | 5 |
| 2025 | Advances in artificial intelligence for diabetes prediction: insights from a systematic literature reviewabstractDiabetes mellitus (DM), a prevalent metabolic disorder, has significant global health implications. The advent of machine learning (ML) has revolutionized the ability to predict and manage diabetes early, offering new avenues to mitigate its impact. This systematic review examined 53 articles on ML applications for diabetes prediction, focusing on datasets, algorithms, training methods, and evaluation metrics. Various datasets, such as the Singapore National Diabetic Retinopathy Screening Program, REPLACE-BG, National Health and Nutrition Examination Survey (NHANES), and Pima Indians Diabetes Database (PIDD), have been explored, highlighting their unique features and challenges, such as class imbalance. This review assesses the performance of various ML algorithms, such as Convolutional Neural Networks (CNN), Support Vector Machines (SVM), Logistic Regression, and XGBoost, for the prediction of diabetes outcomes from multiple datasets. In addition, it explores explainable AI (XAI) methods such as Grad-CAM, SHAP, and LIME, which improve the transparency and clinical interpretability of AI models in assessing diabetes risk and detecting diabetic retinopathy. Techniques such as cross-validation, data augmentation, and feature selection are discussed in terms of their influence on the versatility and robustness of the model. Some evaluation techniques involving k-fold cross-validation, external validation, and performance indicators such as accuracy, area under curve, sensitivity, and specificity are presented. The findings highlight the usefulness of ML in addressing the challenges of diabetes prediction, the value of sourcing different data types, the need to make models explainable, and the need to keep models clinically relevant. This study highlights significant implications for healthcare professionals, policymakers, technology developers, patients, and researchers, advocating interdisciplinary collaboration and ethical considerations when implementing ML-based diabetes prediction models. By consolidating existing knowledge, this SLR outlines future research directions aimed at improving diagnostic accuracy, patient care, and healthcare efficiency through advanced ML applications. This comprehensive review contributes to the ongoing efforts to utilize artificial intelligence technology for a better prediction of diabetes, ultimately aiming to reduce the global burden of this widespread disease. • Systematic Literature Review of 53 articles on ML in diabetes prediction, exploring datasets, methods and evaluation benchmarks. • SVM, XGBoost, Logistic Regression, CNN are evaluated for diabetes prediction. • Accuracy, sensitivity, specificity, precission are used as metrics; validation strategies are outlined. • Finds the need for collaboration between tech-experts, doctors, patients and policymakers in ML healthcare use. • Recommends future research into more advanced ML algorithms focused on explainability for diabetes prediction and healthcare delivery. Pir Bakhsh Khokhar, Carmine Gravino, Fabio Palomba |
Artif. Intell. Medicine | 2 |
| 2025 | Beyond domain dependency in security requirements identificationabstractEarly security requirements identification is crucial in software development, facilitating the integration of security measures into IT networks and reducing time and costs throughout software life-cycle. This paper addresses the limitations of existing methods that leverage Natural Language Processing (NLP) and machine learning techniques for detecting security requirements. These methods often fall short in capturing syntactic and semantic relationships, face challenges in adapting across domains, and rely heavily on extensive domain-specific data. In this paper we focus on identifying the most effective approaches for this task, highlighting both domain-specific and domain-independent strategies. Our methodology encompasses two primary streams of investigation. First, we explore shallow machine learning techniques, leveraging word embeddings. We test ensemble methods and grid search within and across domains, evaluating on three industrial datasets. Next, we develop several domain-independent models based on BERT, tailored to better detect security requirements by incorporating data on software weaknesses and vulnerabilities. Our findings reveal that ensemble and grid search methods prove effective in domain-specific and domain-independent experiments, respectively. However, our custom BERT models showcase domain independence and adaptability. Notably, the CweCveCodeBERT model excels in Precision and F1-score, outperforming existing approaches significantly. It improves F1-score by ∼ 3% and Precision by ∼ 14% over the best approach currently in the literature. BERT-based models, especially with specialized pre-training, show promise for automating security requirement detection. This establishes a foundation for software engineering researchers and practitioners to utilize advanced NLP to improve security in early development phases, fostering the adoption of these state-of-the-art methods in real-world scenarios. Francesco Casillo, Vincenzo Deufemia, Carmine Gravino |
Inf. Softw. Technol. | 3 |
| 2025 | LLM-Based Automation of COSMIC Functional Size Measurement From Use CasesabstractCOmmon Software Measurement International Consortium (COSMIC) Functional Size Measurement is a method widely used in the software industry to quantify user functionality and measure software size, which is crucial for estimating development effort, cost, and resource allocation. COSMIC measurement is a manual task that requires qualified professionals and effort. To support professionals in COSMIC measurement, we propose an automatic approach, CosMet, that leverages Large Language Models to measure software size starting from use cases specified in natural language. To evaluate the proposed approach, we developed a web tool that implements CosMet using GPT-4 and conducted two studies to assess the approach quantitatively and qualitatively. Initially, we experimented with CosMet on seven software systems, encompassing 123 use cases, and compared the generated results with the ground truth created by two certified professionals. Then, seven professional measurers evaluated the analysis achieved by CosMet and the extent to which the approach reduces the measurement time. The first study's results revealed that CosMet is highly effective in analyzing and measuring use cases. The second study highlighted that CosMet offers a transparent and interpretable analysis, allowing practitioners to understand how the measurement is derived and make necessary adjustments. Additionally, it reduces the manual measurement time by 60-80%. Gabriele De Vito, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Fabio Palomba |
IEEE Trans. Software Eng. | 4 |
| 2025 | RECOVER: Toward Requirements Generation From Stakeholders' ConversationsabstractStakeholders’ conversations in requirements elicitation meetings hold valuable insights into system and client needs. However, manually extracting requirements is time-consuming, labor-intensive, and prone to errors and biases. While current state-of-the-art methods assist in summarizing stakeholder conversations and classifying requirements based on their nature, there is a noticeable lack of approaches capable of both identifying requirements within these conversations and generating corresponding system requirements. These approaches would assist requirement identification, reducing engineers’ workload, time, and effort. They would also enhance accuracy and consistency in documentation, providing a reliable foundation for further analysis. To address this gap, this paper introduces RECOVER (Requirements EliCitation frOm conVERsations), a novel conversational requirements engineering approach that leverages natural language processing and large language models (LLMs) to support practitioners in automatically extracting system requirements from stakeholder interactions by analyzing individual conversation turns. The approach is evaluated using a mixed-method research design that combines statistical performance analysis with a user study involving requirements engineers, targeting two levels of granularity. First, at the conversation turn level, the evaluation measures RECOVER’s accuracy in identifying requirements-relevant dialogue and the quality of generated requirements in terms of correctness, completeness, and actionability. Second, at the entire conversation level, the evaluation assesses the overall usefulness and effectiveness of RECOVER in synthesizing comprehensive system requirements from full stakeholder discussions. Empirical evaluation of RECOVER shows promising performance, with generated requirements demonstrating satisfactory correctness, completeness, and actionability. The results also highlight the potential of automating requirements elicitation from conversations as an aid that enhances efficiency while maintaining human oversight. Gianmario Voria, Francesco Casillo, Carmine Gravino, Gemma Catolino, Fabio Palomba |
IEEE Trans. Software Eng. | 3 |
| 2024 | ReFAIR: Toward a Context-Aware Recommender for Fairness Requirements EngineeringabstractMachine learning (ML) is increasingly being used as a key component of most software systems, yet serious concerns have been raised about the fairness of ML predictions. Researchers have been proposing novel methods to support the development of fair machine learning solutions. Nonetheless, most of them can only be used in late development stages, e.g., during model training, while there is a lack of methods that may provide practitioners with early fairness analytics enabling the treatment of fairness throughout the development lifecycle. This paper proposes ReFair, a novel context-aware requirements engineering framework that allows to classify sensitive features from User Stories. By exploiting natural language processing and word embedding techniques, our framework first identifies both the use case domain and the machine learning task to be performed in the system being developed; afterward, it recommends which are the context-specific sensitive features to be considered during the implementation. We assess the capabilities of ReFair by experimenting it against a synthetic dataset---which we built as part of our research---composed of 12,401 User Stories related to 34 application domains. Our findings showcase the high accuracy of ReFair, other than highlighting its current limitations. Carmine Ferrara, Francesco Casillo, Carmine Gravino, Andrea De Lucia, Fabio Palomba |
ICSE | 3 |
| 2024 | On the adoption and effects of source code reuse on defect proneness and maintenance effortabstractAbstract Software reusability mechanisms, like inheritance and delegation in Object-Oriented programming, are widely recognized as key instruments of software design that reduce the risks of source code being affected by defects, other than to reduce the effort required to maintain and evolve source code. Previous work has traditionally employed source code reuse metrics for prediction purposes, e.g., in the context of defect prediction. However, our research identifies two noticeable limitations of the current literature. First, still little is known about the extent to which developers actually employ code reuse mechanisms over time. Second, it is still unclear how these mechanisms may contribute to explaining defect-proneness and mainten0ance effort during software evolution. We aim at bridging this gap of knowledge, as an improved understanding of these aspects might provide insights into the actual support provided by these mechanisms, e.g., by suggesting whether and how to use them for prediction purposes. We propose an exploratory study, conducted on 12Javaprojects–over 44,900 commits–of theDefects4Jdataset, aiming at (1) assessing how developers use inheritance and delegation during software evolution; and (2) statistically analyzing the impact of inheritance and delegation on fault proneness and maintenance effort. Our results let emerge various usage patterns that describe the way inheritance and delegation vary over time. In addition, we find out that inheritance and delegation are statistically significant factors that influence both source code defect-proneness and maintenance effort. Giammaria Giordano, Gerardo Festa, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, Carmine Gravino |
Empir. Softw. Eng. | 6 |
| 2024 | SENEM: A software engineering-enabled educational metaverseabstractThe term metaverse refers to a persistent, virtual, three-dimensional environment where individuals may communicate, engage, and collaborate. One of the most multifaceted and challenging use cases of the metaverse is education, where educators and learners may require multiple technical, social, psychological, and interaction instruments to accomplish their learning objectives. While the characteristics of the metaverse might nicely fit the problem’s needs, our research points out a noticeable lack of knowledge into (1) the specific requirements that an educational metaverse should actually fulfill to let educators and learners successfully interact towards their objectives and (2) how to design an appropriate educational metaverse for both educators and learners. In this paper, we aim to bridge this knowledge gap by proposing SENEM, a novel software engineering-enabled educational metaverse. We first elicit a set of functional requirements that an educational metaverse should fulfill. In this respect, we conduct a literature survey to extract the currently available knowledge on the matter discussed by the research community, and afterward, we assess and complement such knowledge through semi-structured interviews with educators and learners. Upon completing the requirements elicitation stage, we then build our prototype implementation of SENEM, a metaverse that makes available to educators and learners the features identified in the previous stage. Finally, we evaluate the tool in terms of learnability, efficiency, and satisfaction through a Rapid Iterative Testing and Evaluation research approach, leading us to the iterative refinement of our prototype. Through our survey strategy, we extracted nine requirements that guided the tool development that the study participants positively evaluated. Our study reveals that the target audience appreciates the elicited design strategy. Our work has the potential to form a solid contribution that other researchers can use as a basis for further improvements. Viviana Pentangelo, Dario Di Dario, Stefano Lambiase, Filomena Ferrucci, Carmine Gravino, Fabio Palomba |
Inf. Softw. Technol. | 5 |
| 2023 | Toward a Secure Educational Metaverse: A Tale of Blockchain Design for Educational EnvironmentsabstractIn the era of social distancing, distance learning represents a crucial educational challenge. Several 2D information technologies have been provided, yet these share multiple limitations and have negative social, educational, and psychological implications for learners. Metaverse promises to revolutionize education as we know it: this is a persistent, virtual, three-dimensional environment that is supposed to address most of the limitations of 2D information technologies. Nonetheless, there are still software engineering challenges to face to enable such a metaverse, especially when turning to software security and privacy. In this paper, we aim at performing the first steps toward an improved understanding of the security perspective of educational metaverse, by analyzing how blockchain can be employed within educational environments and how applications may be designed. Our ultimate goal is to provide insights into how blockchain can be further tailored in the context of educational metaverse. We conduct a systematic literature review, which targets 20 primary studies. The key findings of the study showcase the use of blockchain in 3 educational tasks, other than describing the blockchain design approaches, which protocol they commonly use and the associated limitations. We conclude by developing a conceptualization of a blockchain-based educational metaverse. Dario Di Dario, Umberto Bilotti, Maurizio Sibilio, Carmine Gravino, Fabio Palomba |
SEAA | 4 |
| 2023 | ECHO: An Approach to Enhance Use Case Quality Exploiting Large Language ModelsabstractUML use cases are commonly used in software engineering to specify the functional requirements of a system since they are an effective tool for interacting with stakeholders thanks to the use of natural languages. However, producing high-quality use cases can be challenging due to the lack of precise guidelines and suitable tools. This can lead to problems, e.g. inaccuracy and incompleteness, in the derived software artifacts and the final product. Recent advancements in Natural Language Processing and Large Language Models (LLMs) can provide the premises for developing tools supporting activities based on natural languages. In this paper, we propose ECHO, a novel approach for supporting software engineers in enhancing the quality of UML use cases using LLMs. Our approach consists of a co-prompt engineering approach and an iterative and interactive process with the LLM to improve the quality of use cases, based on practitioners’ feedback. To prove the feasibility of the proposal, we instantiated the approach using ChatGPT and performed a controlled experiment to assess its effectiveness by involving seven software engineering professionals. Three were part of the experimental group and used ECHO to improve the quality of the use cases. Three others were the control group and enhanced the quality of use cases manually. Finally, the last participant acted as an oracle, blind w.r.t. the groups, and evaluated the quality of the enhanced use cases, both qualitatively by means of a questionnaire, and quantitatively, by means of the Use Case Points metric. Results show that ECHO can effectively support software engineers to improve use cases’ quality thanks to the prompts suitably designed to interact with ChatGPT. Gabriele De Vito, Fabio Palomba, Carmine Gravino, Sergio Di Martino, Filomena Ferrucci |
SEAA | 3 |
| 2022 | The Impact of Parameters Optimization in Software Prediction ModelsabstractSeveral studies have raised concerns about the performance of estimation techniques if employed with default parameters provided by specific development toolkits, e.g., Weka. In this paper, we evaluate the impact of parameter optimization with nine different estimation techniques in the Software Development Effort Estimation (SDEE) and Software Fault Prediction (SFP) domains to provide more generic findings of the impact of parameter optimization. To this aim, we employ three datasets from the domain of SDEE (China, Maxwell, Nasa) and three different regression-based datasets from the SFP domain (Ant, Xalan, Xerces). Regarding parameter optimization, we consider four optimization algorithms from different families: Grid Search and Random Search, Simulated Annealing, and Bayesian Optimization. The estimation techniques are: Support Vector Machine, Random Forest, Classification and Regression Tree, Neural Networks, Averaged Neural Networks, k-Nearest Neighbor, Partial Least Square, MultiLayer Perceptron, and Gradient Boosting Machine. Results reveal that, with both SDEE and SFP datasets, seven out of nine estimation techniques require optimization/configuration of at least one parameter. In majority of the cases, the parameters of the employed estimation techniques are sensitive to the optimization of specific types of data. Moreover, not all the parameters need to be optimized as some of them are not sensitive to optimization. Asad Ali 0006, Carmine Gravino |
SEAA | 2 |
| 2022 | Using COSMIC to measure functional size of software: a Systematic Literature ReviewabstractWe present a systematic literature review performed to understand and summarize the application of COSMIC which is a functional size measurement (FSM) method, mainly applied for estimating software development effort. The results reveals that it is considered to be suitable for a broader range of application domains, e.g., Web applications, Mobile app. Furthermore, the review shows that a lot has been done also for automating the use of COSMIC as well as for rapidly and early applying the method through some approximations. Vincenzo Luigi Martino, Carmine Gravino |
SEAA | 2 |
| 2022 | On the Evolution of Inheritance and Delegation Mechanisms and Their Impact on Code QualityabstractSource code reuse is considered one of the holy grails of modern software development. Indeed, it has been widely demonstrated that this activity decreases software development and maintenance costs while increasing its overall trustwor-thiness. The Object-Oriented (OO) paradigm provides different internal mechanisms to favor code reuse, i.e., specification inheritance, implementation inheritance, and delegation. While previous studies investigated how inheritance relations impact source code quality, there is still a lack of understanding of their evolutionary aspects and, more particular, of how these mechanisms may impact source code quality over time. To bridge this gap of knowledge, this paper proposes an empirical investigation into the evolution of specification inheritance, implementation inheritance, and delegation and their impact on the variability of source code quality attributes. First, we assess how the implementation of those mechanisms varies over 15 releases of three software systems. Second, we devise a statistical approach with the aim of understanding how inheritance and delegation let source code quality—as indicated by the severity of code smells—vary in either positive or negative manner. The key results of the study indicate that inheritance and delegation evolve over time, but not in a statistically significant manner. At the same time, their evolution often leads code smell severity to be reduced, hence possibly contributing to improve code maintainability. Giammaria Giordano, Antonio Fasulo, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, Carmine Gravino |
SANER | 6 |
| 2022 | Detecting privacy requirements from User Stories with NLP transfer learning models
Francesco Casillo, Vincenzo Deufemia, Carmine Gravino |
Inf. Softw. Technol. | 3 |
| 2022 | Evaluating the impact of feature selection consistency in software prediction
Asad Ali 0006, Carmine Gravino |
Sci. Comput. Program. | 2 |
| 2022 | The Importance of the Correlation in Crossover ExperimentsabstractContext:In empirical software engineering, crossover designs are popular for experiments comparing software engineering techniques that must be undertaken by human participants. However, their value depends on the correlation ($r$) between the outcome measures on the same participants. Software engineering theory emphasizes the importance of individual skill differences, so we would expect the values of$r$to be relatively high. However, few researchers have reported the values of$r$.Goal:To investigate the values of$r$found in software engineering experiments.Method:We undertook simulation studies to investigate the theoretical and empirical properties of$r$. Then we investigated the values of$r$observed in 35 software engineering crossover experiments.Results:The level of$r$obtained by analysing our 35 crossover experiments was small. Estimates based on means, medians, and random effect analysis disagreed but were all between 0.2 and 0.3. As expected, our analyses found large variability among the individual$r$estimates for small sample sizes, but no indication that$r$estimates were larger for the experiments with larger sample sizes that exhibited smaller variability.Conclusions:Low observed$r$values cast doubts on the validity of crossover designs for software engineering experiments. However, if the cause of low$r$values relates to training limitations or toy tasks, this affectsallSoftware Engineering (SE) experiments involving human participants. For all human-intensive SE experiments, we recommend more intensive training and then tracking the improvement of participants as they practice using specific techniques, before formally testing the effectiveness of the techniques. Barbara A. Kitchenham, Lech Madeyski, Giuseppe Scanniello, Carmine Gravino |
IEEE Trans. Software Eng. | 4 |
| 2021 | Combining CNN with DS3 for Detecting Bug-prone Modules in Cross-version ProjectsabstractThe paper focuses on Cross-Version Defect Prediction (CVDP) where the classification model is trained on information of the prior version and then tested to predict defects in the components of the last release. To avoid the distribution differences which could negatively impact the performances of machine learning based model, we consider Dissimilarity-based Sparse Subset Selection (DS3) technique for selecting meaningful representatives to be included in the training set. Furthermore, we employ a Convolutional Neural Network (CNN) to generate structural and semantic features to be merged with the traditional software measures to obtain a more comprehensive list of predictors. To evaluate the usefulness of our proposal for the CVDP scenario, we perform an empirical study on a total of 20 cross-version pairs from 10 different software projects. To build prediction models we consider Logistic Regression (LR) and Random Forest (RF) and we adopt 3 evaluation criteria (i.e., F-measure, G-mean, Balance) to assess the prediction accuracy. Our results show that the use of CNN with both LR and RF models has a significant impact, with an improvement of ∼20% for each evaluation criteria. Differently, we notice that DS3does not impact significantly in improving prediction accuracy. Andrea Fiore, Alfonso Russo, Carmine Gravino, Michele Risi |
SEAA | 3 |
| 2021 | Using the Normalized Levenshtein Distance to Analyze Relationship between Faults and Local Variables with Confusing Names: A further Investigation (S)abstractThis paper exploits further uses of NLD (Normalized Levenshtein Distance), proposed in a recent study, to quantify the level of confusion of variables with the aim of verifying if they can provide indications about the presence of faults.We provide further evidence that fault prediction models based on the considered NLD measures can provide accurate estimations. Carmine Gravino, Alessandra Orsi, Michele Risi |
SEKE | 1 |
| 2021 | Improving software effort estimation using bio-inspired algorithms to select relevant features: An empirical study
Asad Ali 0006, Carmine Gravino |
Sci. Comput. Program. | 2 |
| 2021 | An empirical comparison of validation methods for software prediction modelsabstractAbstract Model validation methods (e.g., k‐fold cross‐validation) use historical data to predict how well an estimation technique (e.g., random forest) performs on the current (or future) data. Studies in the contexts of software development effort estimation (SDEE) and software fault prediction (SFP) have used and investigated different model validation methods. However, no conclusive indications to suggest which model validation method has a major impact on the prediction accuracy and stability of estimation techniques. Some studies have investigated model validation methods specific to data about either SDEE or SFP. To the best of our knowledge, there is no study in the literature, which has employed different validation methods both with SDEE and SFP data. The aim of this paper is to consider different methods (10) from the family of cross‐validation (CV) and bootstrap validation methods to identify which one contributes to obtaining a better prediction accuracy for both types of data. We also evaluate which model validation methods allow the estimation techniques to provide stable performances (i.e., with lower variance). To this aim, we present an empirical study involving six datasets from the domain of SDEE and six datasets from the SFP domain. The results reveal that repeated 10‐fold CV with SDEE and optimistic boot with SFP data are the model validation methods that provide a better prediction accuracy in a greater number of experiments than the other model validation methods. Furthermore, a model validation method can improve the prediction accuracy up to 60% with SDEE data and up to 36% when employing SFP data. The analysis also reveals that repeated fivefold CV produces more stable performances when the experiments are repeated on the same data. Asad Ali 0006, Carmine Gravino |
J. Softw. Evol. Process. | 2 |
| 2020 | Assessing the effectiveness of approximate functional sizing approaches for effort estimation
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro |
Inf. Softw. Technol. | 3 |
| 2020 | Design and automation of a COSMIC measurement procedure based on UML models
Gabriele De Vito, Filomena Ferrucci, Carmine Gravino |
Softw. Syst. Model. | 3 |
| 2019 | Using Bio-Inspired Features Selection Algorithms in Software Effort Estimation: A Systematic Literature ReviewabstractFeature selection algorithms select the best and relevant set of features of the datasets which leads to an increase in the accuracy of predictions when employed with the machine learning techniques. Different feature selection algorithms are used in the domain of Software Development Effort Estimations (SDEE) and recently the use of bio-inspired feature selection algorithms got the attention of the researchers, which provided the best results in terms of the accuracy measures. In this paper, we manage to systematically evaluate and assess different bio-inspired feature selection algorithms which have been employed and investigated in the studies related to SDEE with the aim of increasing the accuracy of estimations. To the best of our knowledge, there is no Systematic Literature Review (SLR) which investigated the use of bio-inspired algorithms in SDEE. Since, the use of bio-inspired algorithms in the area of SDEE started in the late 2000, we have considered the studies published between 2007-2018. We have selected about 30 different studies from five digital libraries, i.e., IEEE explore, Springer, ScienceDirect, ACM digital library, and Google Scholar, after the filtering of inclusion/exclusion and quality assessment criteria. The main findings of our SLR are that Genetic Algorithms (GA) and Particle Swarm Optimizations (PSO) are widely used bio-inspired algorithms. Moreover, GA and PSO are the algorithms which outperform baseline estimation techniques (estimation techniques employed without any feature selection algorithms) in more number of experiments, in terms of prediction accuracy. Asad Ali 0006, Carmine Gravino |
SEAA | 2 |
| 2019 | Can Expert Opinion Improve Effort Predictions When Exploiting Cross-Company Datasets? - A Case Study in a Small/Medium Company
Filomena Ferrucci, Carmine Gravino |
PROFES | 2 |
| 2019 | A systematic literature review of software effort prediction using machine learning methodsabstractAbstract Machine learning (ML) techniques have been widely investigated for building prediction models, able to estimate software development effort as well as to improve the accuracy of other estimation techniques. The objective of this paper is to systematically review the recent studies which used and discussed the software effort estimation models built using ML techniques. The performed literature review is based on the empirical studies published in the time period of January 1991 to December 2017, by employing widely used guidelines. The review has selected a total of 75 primary studies after the careful filtering of inclusion/exclusion and quality assessment criteria. The performed analysis reveals that artificial neural network (ANN) as ML model, NASA as dataset, and mean magnitude of relative error (MMRE) as accuracy measure are widely used in the selected studies. ANN and support vector machine (SVM) are the two techniques which have outperformed other ML techniques in more studies. Regression techniques are the mostly used among the non‐ML techniques, which outperformed other ML techniques in about 19 studies. Moreover, SVM and regression techniques in combination are characterized by better predictions when compared with other ML and non‐ML techniques. Asad Ali 0006, Carmine Gravino |
J. Softw. Evol. Process. | 2 |
| 2018 | Impact of Design Pattern Implementation Variants on the Retrieval Effectiveness of a Recovery Tool: An Exploratory StudyabstractThis paper investigates howimplementation variantsof design patterns impact on the retrieval effectiveness of a design pattern recovery tool. Specifically, we first defined several implementation variants of Adapter and Observer design patterns, by introducing constraints or relaxations on their canonical form. Then, we analyze the relationship between the complexity of these definitions and the precision and time needed by a design pattern recovery process we proposed in the past. To this end, we apply ePAD, an Eclipse plug-in for design pattern recovery, to eight software systems. We show that there exist interesting issues about the relationship between the complexity of the defined variants and the precision and time needed to recover their instances. Andrea De Lucia, Vincenzo Deufemia, Carmine Gravino, Michele Risi |
SEAA | 3 |
| 2018 | Do software models based on the UML aid in source-code comprehensibility? Aggregating evidence from 12 controlled experiments
Giuseppe Scanniello, Carmine Gravino, Marcela Genero, José A. Cruz-Lemus, Genny Tortora, Michele Risi, Gabriella Dodero |
Empir. Softw. Eng. | 2 |
| 2018 | Definition and evaluation of a COSMIC measurement procedure for sizing Web applications in a model-driven development environment
Silvia Abrahão, Lucia De Marco, Filomena Ferrucci, Jaime Gómez, Carmine Gravino, Federica Sarro |
Inf. Softw. Technol. | 5 |
| 2018 | Guest Editorial of Special Section on Software Engineering and Advanced Applications in Information Technology for Software-Intensive Systems
Carmine Gravino, Martin Höst |
Inf. Softw. Technol. | 1 |
| 2018 | Detecting the Behavior of Design Patterns through Model Checking and Dynamic AnalysisabstractWe present a method and tool (ePAD) for the detection of design pattern instances in source code. The approach combines static analysis, based on visual language parsing and model checking, and dynamic analysis, based on source code instrumentation. Visual language parsing and static source code analysis identify candidate instances satisfying the structural properties of design patterns. Successively, model checking statically verifies the behavioral aspects of the candidates recovered in the previous phase. We encode the sequence of messages characterizing the correct behaviour of a pattern as Linear Temporal Logic (LTL) formulae and the sequence diagram representing the possible interaction traces among the objects involved in the candidates as Promela specifications. The model checker SPIN verifies that candidates satisfy the LTL formulae. Dynamic analysis is then performed on the obtained candidates by instrumenting the source code and monitoring those instances at runtime through the execution of test cases automatically generated using a search-based approach. The effectiveness of ePAD has been evaluated by detecting instances of 12 creational and behavioral patterns from six publicly available systems. The results reveal that ePAD outperforms other approaches by recovering more actual instances. Furthermore, on average ePAD achieves better results in terms of correctness and completeness. Andrea De Lucia, Vincenzo Deufemia, Carmine Gravino, Michele Risi |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2017 | How the Use of Design Patterns Affects the Quality of Software Systems: A Preliminary InvestigationabstractIn this paper we analyze at the class level the quality of the software portions including classes participating in design patterns instances (DP classes) with respect to the remaining software portions (NoDP classes). The performed study is based on 10 software systems from which information about design pattern instances and CK (Chidamber and Kemerer) metrics were obtained by exploiting repositories of pattern instances and the tool Understand, respectively. The analysis revealed that the use of design patterns impacts on the quality of the software. Carmine Gravino, Michele Risi |
SEAA | 1 |
| 2017 | A study on the statistical convertibility of IFPUG Function Point, COSMIC Function Point and Simple Function Point
Abedallah Zaid Abualkishik, Filomena Ferrucci, Carmine Gravino, Luigi Lavazza, Roberto Meli, Gabriela Robiolo |
Inf. Softw. Technol. | 3 |
| 2016 | Web Effort Estimation: Function Point Analysis vs. COSMIC
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro |
Inf. Softw. Technol. | 3 |
| 2015 | Towards automating dynamic analysis for behavioral design pattern detectionabstractThe detection of behavioral design patterns is more accurate when a dynamic analysis is performed on the candidate instances identified statically. Such a dynamic analysis requires the monitoring of the candidate instances at run-time through the execution of a set of test cases. However, the definition of such test cases is a time-consuming task if performed manually, even more, when the number of candidate instances is high and they include many false positives. In this paper we present the results of an empirical study aiming at assessing the effectiveness of dynamic analysis based on automatically generated test cases in behavioral design pattern detection. The study considered three behavioral design patterns, namely State, Strategy, and Observer, and three publicly available software systems, namely JHotDraw 5.1, QuickUML 2001, and MapperXML 1.9.7. The results show that dynamic analysis based on automatically generated test cases improves the precision of design pattern detection tools based on static analysis only. As expected, this improvement in precision is achieved at the expenses of recall, so we also compared the results achieved with automatically generated test cases with the more expensive but also more accurate results achieved with manually built test cases. The results of this analysis allowed us to highlight costs and benefits of automating dynamic analysis for design pattern detection. Andrea De Lucia, Vincenzo Deufemia, Carmine Gravino, Michele Risi |
ICSME | 3 |
| 2015 | ePadEvo: A tool for the detection of behavioral design patternsabstractIn this demonstration we present ePADevo, an Eclipse plug-in for recovering design pattern instances from object-oriented source code. The tool is able to recover design pattern instances through a static analysis performed on a data model extracted from source code, and a dynamic analysis performed through the instrumentation and the monitoring of the software system. Dynamic analysis is performed with automatically generated test cases exploiting the EvoSuite tool. Andrea De Lucia, Vincenzo Deufemia, Carmine Gravino, Michele Risi, Ciro Pirolli |
ICSME | 3 |
| 2015 | From Function Points to COSMIC - A Transfer Learning Approach for Effort Estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro |
PROFES | 4 |
| 2015 | Investigating Functional and Code Size Measures for Mobile Applications: A Replicated Study
Filomena Ferrucci, Carmine Gravino, Pasquale Salza, Federica Sarro |
PROFES | 2 |
| 2015 | Studying the Effect of UML-Based Models on Source-Code Comprehensibility: Results from a Long-Term Investigation
Giuseppe Scanniello, Carmine Gravino, Genny Tortora, Marcela Genero, Michele Risi, José A. Cruz-Lemus, Gabriella Dodero |
PROFES | 2 |
| 2015 | On the Effect of Exploiting GPUs for a More Eco-Sustainable Lease of LifeabstractIt has been estimated that about 2% of global carbon dioxide emissions can be attributed to IT systems. Green (or sustainable) computing refers to supporting business critical computing needs with the least possible amount of power. This phenomenon changes the priorities in the design of new software systems and in the way companies handle existing ones. In this paper, we present the results of a research project aimed to develop a migration strategy to give an existing software system a new and more eco-sustainable lease of life. We applied a strategy for migrating a subject system that performs intensive and massive computation to a target architecture based on a Graphics Processing Unit (GPU). We validated our solution on a system for path finding robot simulations. An analysis on execution time and energy consumption indicated that: (i) the execution time of the migrated system is less than the execution time of the original system; and (ii) the migrated system reduces energy waste, so suggesting that it is more eco-sustainable than its original version. Our findings improve the body of knowledge on the effect of using the GPU in green computing. Giuseppe Scanniello, Ugo Erra, Giuseppe Caggianese, Carmine Gravino |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2015 | A fine-grained analysis of the support provided by UML class diagrams and ER diagrams during data model maintenance
Gabriele Bavota, Carmine Gravino, Rocco Oliveto, Andrea De Lucia, Genny Tortora, Marcela Genero, José A. Cruz-Lemus |
Softw. Syst. Model. | 2 |
| 2015 | Documenting Design-Pattern Instances: A Family of Experiments on Source-Code ComprehensibilityabstractDesign patterns are recognized as a means to improve software maintenance by furnishing an explicit specification of class and object interactions and their underlying intent [Gamma et al. 1995]. Only a few empirical investigations have been conducted to assess whether the kind of documentation for design patterns implemented in source code affects its comprehensibility. To investigate this aspect, we conducted a family of four controlled experiments with 88 participants having different experience (i.e., professionals and Bachelor, Master, and PhD students). In each experiment, the participants were divided into three groups and asked to comprehend a nontrivial chunk of an open-source software system. Depending on the group, each participant was, or was not, provided with graphical or textual representations of the design patterns implemented within the source code. We graphically documented design-pattern instances with UML class diagrams. Textually documented instances are directly reported source code as comments. Our results indicate that documenting design-pattern instances yields an improvement in correctness of understanding source code for those participants with an adequate level of experience. Giuseppe Scanniello, Carmine Gravino, Michele Risi, Genny Tortora, Gabriella Dodero |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2014 | On the impact of UML analysis models on source-code comprehensibility and modifiabilityabstractWe carried out a family of experiments to investigate whether the use of UML models produced in the requirements analysis process helps in the comprehensibility and modifiability of source code. The family consists of a controlled experiment and 3 external replications carried out with students and professionals from Italy and Spain. 86 participants with different abilities and levels of experience with UML took part. The results of the experiments were integrated through the use of meta-analysis. The results of both the individual experiments and meta-analysis indicate that UML models produced in the requirements analysis process influence neither the comprehensibility of source code nor its modifiability. Giuseppe Scanniello, Carmine Gravino, Marcela Genero, José A. Cruz-Lemus, Genny Tortora |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2013 | Class level fault prediction using software clusteringabstractDefect prediction approaches use software metrics and fault data to learn which software properties associate with faults in classes. Existing techniques predict fault-prone classes in the same release (intra) or in a subsequent releases (inter) of a subject software system. We propose an intra-release fault prediction technique, which learns from clusters of related classes, rather than from the entire system. Classes are clustered using structural information and fault prediction models are built using the properties of the classes in each cluster. We present an empirical investigation on data from 29 releases of eight open source software systems from the PROMISE repository, with predictors built using multivariate linear regression. The results indicate that the prediction models built on clusters outperform those built on all the classes of the system. Giuseppe Scanniello, Carmine Gravino, Andrian Marcus, Tim Menzies |
ASE | 2 |
| 2013 | Using tabu search to configure support vector regression for effort estimationabstractRecent studies have reported that Support Vector Regression (SVR) has the potential as a technique for software development effort estimation. However, its prediction accuracy is heavily influenced by the setting of parameters that needs to be done when employing it. No general guidelines are available to select these parameters, whose choice also depends on the characteristics of the dataset being used. This motivated the work described in (Corazza et al. 2010 ), extended herein. In order to automatically select suitable SVR parameters we proposed an approach based on the use of the meta-heuristics Tabu Search (TS). We designed TS to search for the parameters of both the support vector algorithm and of the employed kernel function, namely RBF. We empirically assessed the effectiveness of the approach using different types of datasets (single and cross-company datasets, Web and not Web projects) from the PROMISE repository and from the Tukutuku database. A total of 21 datasets were employed to perform a 10-fold or a leave-one-out cross-validation, depending on the size of the dataset. Several benchmarks were taken into account to assess both the effectiveness of TS to set SVR parameters and the prediction accuracy of the proposed approach with respect to widely used effort estimation techniques. The use of TS allowed us to automatically obtain suitable parameters’ choices required to run SVR. Moreover, the combination of TS and SVR significantly outperformed all the other techniques. The proposed approach represents a suitable technique for software development effort estimation. Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro, Emilia Mendes |
Empir. Softw. Eng. | 4 |
| 2013 | Assessing the Effectiveness of Sequence Diagrams in the Comprehension of Functional Requirements: Results from a Family of Five ExperimentsabstractModeling is a fundamental activity within the requirements engineering process and concerns the construction of abstract descriptions of requirements that are amenable to interpretation and validation. The choice of a modeling technique is critical whenever it is necessary to discuss the interpretation and validation of requirements. This is particularly true in the case of functional requirements and stakeholders with divergent goals and different backgrounds and experience. This paper presents the results of a family of experiments conducted with students and professionals to investigate whether the comprehension of functional requirements is influenced by the use of dynamic models that are represented by means of the UML sequence diagrams. The family contains five experiments performed in different locations and with 112 participants of different abilities and levels of experience with UML. The results show that sequence diagrams improve the comprehension of the modeled functional requirements in the case of high ability and more experienced participants. Silvia Abrahão, Carmine Gravino, Emilio Insfrán, Giuseppe Scanniello, Genny Tortora |
IEEE Trans. Software Eng. | 2 |
| 2012 | Do Professional Developers Benefit from Design Pattern Documentation? A Replication in the Context of Source Code Comprehension
Carmine Gravino, Michele Risi, Giuseppe Scanniello, Genny Tortora |
MoDELS | 1 |
| 2011 | Clustering and lexical information support for the recovery of design pattern in source codeabstractWe propose an approach that leverages lexical information and fuzzy clustering to reduce the number of the design pattern instances that existing approaches based on structural information (i.e., navigating the dependencies among software elements) erroneously recover in source code. To assess the effectiveness of the techniques, we present the results of a case study conducted on four open source software systems implemented in java. The data analysis indicates that the use of lexical information and fuzzy clustering improves the correctness of the results achieved by existing design pattern recovery approaches based on structural information, while preserving the number of design pattern instances correctly identified. Simone Romano 0001, Giuseppe Scanniello, Michele Risi, Carmine Gravino |
ICSM | 4 |
| 2011 | Identifying the Weaknesses of UML Class Diagrams during Data Model Comprehension
Gabriele Bavota, Carmine Gravino, Rocco Oliveto, Andrea De Lucia, Genny Tortora, Marcela Genero, José A. Cruz-Lemus |
MoDELS | 2 |
| 2011 | Using Web Objects for Development Effort Estimation of Web Applications: A Replicated Study
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro |
PROFES | 3 |
| 2011 | A Genetic Algorithm to Configure Support Vector Machines for Predicting Fault-Prone Components
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro |
PROFES | 3 |
| 2011 | How Multi-Objective Genetic Programming Is Effective for Software Development Effort Estimation?
Filomena Ferrucci, Carmine Gravino, Federica Sarro |
SSBSE | 2 |
| 2011 | Investigating the use of Support Vector Regression for web effort estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes |
Empir. Softw. Eng. | 4 |
| 2010 | A controlled experiment for assessing the contribution of design pattern documentation on software maintenanceabstractIn this paper we present the preliminary results of a controlled experiment to assess the contribution provided by the design patterns on the maintenance of source code. In particular, the study aimed at assessing the effort and the efficiency to perform maintenance operations in case design pattern instances are properly documented and provided to the maintainer. The context of the experiment is constituted of Master Students in Computer Science at the University of Basilicata. The preliminary analysis conducted on the gathered data revealed that the effort is significantly reduced in case design pattern instances are properly documented and provided to the subjects. Similarly, the efficiency is significantly better in case the documentation of design pattern instances is used to accomplish maintenance operations. Giuseppe Scanniello, Carmine Gravino, Michele Risi, Genny Tortora |
ESEM | 2 |
| 2010 | An Eclipse plug-in for the detection of design pattern instances through static and dynamic analysisabstractThe extraction of design pattern information from software systems can provide conspicuous insight to software engineers on the software structure and its internal characteristics. In this demonstration we present ePAD, an Eclipse plug-in for recovering design pattern instances from object-oriented source code. The tool is able to recover design pattern instances through a structural analysis performed on a data model extracted from source code, and a behavioral analysis performed through the instrumentation and the monitoring of the software system. ePAD is fully configurable since it allows software engineers to customize the design pattern recovery rules and the layout used for the visualization of the recovered instances. Andrea De Lucia, Vincenzo Deufemia, Carmine Gravino, Michele Risi |
ICSM | 3 |
| 2010 | An experimental comparison of ER and UML class diagrams for data modelling
Andrea De Lucia, Carmine Gravino, Rocco Oliveto, Genny Tortora |
Empir. Softw. Eng. | 2 |
| 2009 | On the effectiveness of dynamic modeling in UML: Results from an external replicationabstractThis paper describes the results of an external replication of an experiment for assessing whether the use of dynamic modeling influences the comprehension of software requirements. The results of the original experiment conducted in Italy did not confirm that there was a significant difference in the comprehension of software requirements when dynamic modeling is used. The goal of the replication was therefore to verify these findings with a group of more experienced students at the Universidad Politeacutecnica de Valencia (UPV) in Spain. The results shows that the use of dynamic modeling does significantly improve the comprehension of software requirements, thus providing evidence that dynamic modeling facilitates the interpretation and comprehension of requirements. Silvia Abrahão, Emilio Insfrán, Carmine Gravino, Giuseppe Scanniello |
ESEM | 3 |
| 2009 | Applying support vector regression for web effort estimation using a cross-company datasetabstractSupport vector regression (SVR) is a new generation of machine learning algorithms, suitable for predictive data modeling problems. The objective of this paper is to investigate the effectiveness of SVR for Web effort estimation, in particular when dealing with a cross-company dataset. To gain a deeper insight on the method, we carried out an empirical study using four kernels for SVR, namely linear, polynomial, Gaussian, and sigmoid. Moreover, we used two variables' preprocessing strategies (normalization and logarithmic), and two different dependent variables (effort and inverse effort). As a result, SVR was applied using six different configurations for each kernel. As for the dataset, we employed the Tukutuku database, which is widely adopted in Web effort estimation studies. A hold-out approach was adopted to evaluate the prediction accuracy for all the configurations, using two training sets, each containing data on 130 projects randomly selected, and two test sets, each containing the remaining 65 projects. As benchmark, SVR-based predictions were also compared to predictions obtained using manual stepwise regression, case-based reasoning, and Bayesian networks. Our results suggest that SVR performed well, since on the first hold-out, the linear kernel with a logarithmic transformation of variables provided significantly superior prediction accuracy than all the other techniques, while for the second hold-out, the Gaussian kernel achieved significantly superior predictions than all other techniques, except for manual stepwise regression. Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes |
ESEM | 4 |
| 2009 | An Empirical Study on the Use of Web-COBRA and Web Objects to Estimate Web Application Development Effort
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino |
ICWE | 3 |
| 2009 | Using Support Vector Regression for Web Development Effort Estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes |
IWSM/Mensura | 4 |
| 2009 | Using Tabu Search to Estimate Software Development Effort
Filomena Ferrucci, Carmine Gravino, Rocco Oliveto, Federica Sarro |
IWSM/Mensura | 2 |
| 2009 | Design pattern recovery through visual language parsing and source code analysis
Andrea De Lucia, Vincenzo Deufemia, Carmine Gravino, Michele Risi |
J. Syst. Softw. | 3 |
| 2009 | Measures and Techniques for Effort Estimation of Web Applications: an Empirical Study Based on a Single-Company Dataset
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes |
J. Web Eng. | 3 |
| 2008 | Data Model Comprehension: An Empirical Comparison of ER and UML Class DiagramsabstractWe present the results of two controlled experiments to compare ER and UML class diagrams, in order to find out which of the models provides better support during the comprehension of data models. The experiment involved Master and Bachelor students performing comprehension tasks on data models represented by ER or UML class diagrams. The achieved results show that UML class diagrams significantly improve the comprehension level achieved by subjects. Moreover, having different subjects with different levels of ability and experience allowed us to also make some considerations on the influence of such factors on the comprehension performances. Andrea De Lucia, Carmine Gravino, Rocco Oliveto, Genny Tortora |
ICPC | 2 |
| 2008 | An Empirical Investigation on Dynamic Modeling in Requirements Engineering
Carmine Gravino, Giuseppe Scanniello, Genny Tortora |
MoDELS | 1 |
| 2008 | Cross-company vs. single-company web effort models using the Tukutuku database: An extended study
Emilia Mendes, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino |
J. Syst. Softw. | 4 |
| 2007 | Comparing Size Measures for Predicting Web Application Development Effort: A Case StudyabstractSize represents one of the most important attribute of software products used to predict software development effort. In the past nine years, several measures have been proposed to estimate the size of Web applications, and it is important to determine which one is most effective to predict Web development effort. To this aim in this paper we report on an empirical analysis where, using data from 15 Web projects developed by a software company, we compare four sets of size measures, using two prediction techniques, namely Forward Stepwise Regression (SWR) and Case-Based Reasoning (CBR). All the measures provided good predictions in terms of MMRE, MdMRE, and Pred(0.25) statistics, for both SWR and CBR. Moreover, when using SWR, length measures and Web Objects gave significant better results than Functional measures, however presented similar results to the Tukutuku measures. As for CBR, results did not show any significant differences amongst the four sets of size measures. Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes |
ESEM | 3 |
| 2007 | A Replicated Study Comparing Web Effort Estimation Techniques
Emilia Mendes, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino |
WISE | 4 |
| 2007 | Effort estimation: how valuable is it for a web company to use a cross-company data set, compared to using its own single-company data set?abstractPrevious studies comparing the prediction accuracy of effort models built using Web cross- and single-company data sets have been inconclusive, and as such replicated studies are necessary to determine under what circumstances a company can place reliance on a cross-company effort model. This paper therefore replicates a previous study by investigating how successful a cross-company effort model is: i) to estimate effort for Web projects that belong to a single company and were not used to build the cross-company model; ii) compared to a single-company effort model. Our single-company data set had data on 15 Web projects from a single company and our cross-company data set had data on 68 Web projects from 25 different companies. The effort estimates used in our analysis were obtained by means of two effort estimation techniques, namely forward stepwise regression and case-based reasoning. Our results were similar to those from the replicated study, showing that predictions based on the single-company model were significantly more accurate than those based on the cross-company model. Emilia Mendes, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino |
WWW | 4 |
| 2006 | Assessing the Usability of a Tool for Developing Adaptive E-learning Processes: an Empirical AnalysisabstractThe correlation between the effort to develop a learning process and early size measures could be used to assess the usability of an employed tool. In particular, when the measures are obtained from the learning process specification and they are relevant effort indicators we can assert that the technical competences of instructional designers are not relevant for the tool usage. We present initial results of applying empirical analysis to confirm a previously usability study performed on the ASCLO-S (Adaptive Self consistent Learning Object SET) editor, a visual language based tool for developing adaptive learning processes. Gennaro Costagliola, Andrea De Lucia, Filomena Ferrucci, Carmine Gravino, Giuseppe Scanniello |
ICALT | 4 |
| 2006 | Effort estimation modeling techniques: a case study for web applicationsabstractA reliable effort estimation is crucial for a successful web application development planning. Several approaches exist to address this issue. Among them, the algorithmic approach is one of the most widely used and investigated methods. It is based on suitable effort prediction models which relate the development effort with project characteristics. The size represents one of the most interesting characteristics of software products and several measures can be defined in order to estimate the size of web systems. Moreover, several techniques have been proposed in the literature to build the effort prediction models. Thus, of special interest should be to establish the most effective size measures to be employed in effort prediction models and the most suitable techniques for the model construction. To this aim some empirical studies have been undertaken so far. Since it is widely recognized that several investigations should be performed to verify/confirm empirical results, in the paper we will report on an empirical analysis we have carried out by exploiting data coming from 15 web projects developed by a software company. In particular, for the analysis we have considered two sets of size measures: Length Measures (e.g. number of pages, number of medias, number of client and server side scripts) and Functional Measures (e.g. external input, external output, external query). Moreover, we have employed different techniques, such as Linear Regression, Regression Tree, and Analogy-Based Estimation, in order to determine the one that provides the best prediction. Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello |
ICWE | 4 |
| 2006 | A COSMIC-FFP Approach to Predict Web Application Development Effort
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello |
J. Web Eng. | 4 |
| 2006 | Constructing Meta-CASE Workbenches by Exploiting Visual Language GeneratorsabstractIn this paper, we propose an approach for the construction of meta-CASE workbenches, which suitably integrates the technology of visual language generation systems, UML metamodeling, and interoperability techniques based on the GXL (graph exchange language) format. The proposed system consists of two major components. Environments for single visual languages are generated by using the modeling language environment generator (MEG), which follows a metamodel/grammar-approach. The abstract syntax of a visual language is defined by UML class diagrams, which serve as a base for the grammar specification of the language. The workbench generator (WoG) allows designers to specify the target workbench by means of a process model given in terms of a suitable activity diagram. Starting from the supplied specification WoG generates the customized workbench by integrating the required environments. Gennaro Costagliola, Vincenzo Deufemia, Filomena Ferrucci, Carmine Gravino |
IEEE Trans. Software Eng. | 4 |
| 2005 | A Cosmic-FFP Approach to Estimate WEB Application Development Effort
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello |
WEBIST | 4 |
| 2005 | Adding symbolic information to picture models: definitions and properties
Gennaro Costagliola, Filomena Ferrucci, Carmine Gravino |
Theor. Comput. Sci. | 3 |
| 2004 | A COSMIC-FFP Based Method to Estimate Web Application Development Effort
Gennaro Costagliola, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello |
ICWE | 3 |
| 2004 | Using COSMIC-FFP for Predicting Web Application Development Effort
Gennaro Costagliola, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello |
SEKE | 3 |
| 2003 | On regular drawn symbolic picture languages
Gennaro Costagliola, Vincenzo Deufemia, Filomena Ferrucci, Carmine Gravino |
Inf. Comput. | 4 |
| 2002 | Using extended positional grammars to develop visual modeling languagesabstractIn this paper we present the approach based on the formalism of Extended Positional Grammars for specifying, designing and implementing visual modeling languages. In order to stress the main characteristics of the approach and highlight its power, we describe the use of the formalism to implement statecharts languages which represent one of the most complex visual modeling languages used in the software engineering field. In the paper special emphasis is put on describing the benefits deriving from the use of such formal specifications such as incrementality, easy customization, and automatic generation of visual programming environments. Such features turn out to be especially important because visual modeling languages are subjected to continuous changes as the history of statecharts languages and UML diagrams shows. Moreover, visual languages can be effectively used only if they are supported by a powerful visual environment within they are embedded and used. Gennaro Costagliola, Vincenzo Deufemia, Filomena Ferrucci, Carmine Gravino |
SEKE | 4 |
| 2001 | Decidability of the consistency problem for regular symbolic picture description languages
Gennaro Costagliola, Vincenzo Deufemia, Filomena Ferrucci, Carmine Gravino |
Inf. Process. Lett. | 4 |