VLDB 2026 Research / reviewers in the wild / expert
Riccardo Rubei
dblp:228/4210
· DBLP profile ↗
25ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-9622-5949ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 21 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Empirical Investigation on the Use of Large Language Models for Performance Bug Detection
Muhammad Imran 0026, Vittorio Cortellessa, Davide Di Ruscio, Riccardo Rubei, Luca Traini |
SANER | 4 |
| 2025 | On the Need for Reproducibility Guidelines for Open-Source Games: A itch.io Case StudyabstractVideo games represent a unique type of complex software project, despite differing in aspects such as methodologies, programming languages, and design patterns. Similar to traditional software, games can be released as open-source software (OSS) projects, thus fostering collaboration and knowledge sharing among developers. However, one of the primary challenges in this domain is the difficulty in finding source code for game development, which is often scattered across various repositories and platforms. Moreover, the lack of proper guidelines to document game projects is still missing, thus worsening this issue. In this paper, we envision a set of initial guidelines leveraging a mining-based methodology considering two different OSS platforms, i.e. itch.io and GitHub. First, we collect data from 765 open-source games from the itch.io platform and map the retrieved games to the corresponding repositories on GitHub, searching for documentation and source code. We further refine the list of games by manually analyzing the repositories, focusing on the quality of the documentation and the presence of source code, ending up with 613 games. On top of this gold set, we elicited a set of ten reproducibility guidelines specifically tailored for games. Our results show that the majority of the games do not have source code available, and the documentation quality is generally low, apart from games with high-rated GitHub projects. In addition, we provide a set of takeaways that can be further investigated by extending the provided guidelines. We believe that our dataset and methodology can be used as a starting point for future research in this domain, providing insights into the challenges and opportunities for assisting newcomers to game development using open-source projects. Claudio Di Sipio, Andrea D'Angelo, Riccardo Rubei, Cristiano Politowski |
CoG | 3 |
| 2025 | Is code coverage of performance tests related to source code features? An empirical study on open-source Java systemsabstractAbstract Performance testing aims to ensure the operational efficiency of software systems. However, many factors influencing the efficacy and adoption of performance tests in practice are not yet fully understood. For instance, while code coverage is widely regarded as a key quality metric for evaluating the efficacy of functional testing suites, there is limited knowledge about the types and levels of coverage that performance tests specifically achieve. Another important factor, often perceived as a barrier to the broader adoption of performance tests yet remaining relatively unexplored, is their extended execution time. In this paper, we examine (i) the coverage of performance testing suites, (ii) the characteristics of source code associated with performance-tested components, and (iii) the time cost of executing performance tests. Our analysis on open-source Java systems reveals that performance tests achieve significantly lower code coverage than functional tests, as expected, and it highlights a significant trade-off between coverage and execution time. Our results also indicate a lack of generalizable characteristics in the source code covered by performance tests. Muhammad Imran 0026, Vittorio Cortellessa, Davide Di Ruscio, Riccardo Rubei, Luca Traini |
Empir. Softw. Eng. | 4 |
| 2025 | Leveraging synthetic trace generation of modeling operations for intelligent modeling assistants using large language modelsabstractContext: Due to the proliferation of generative AI models in different software engineering tasks, the research community has started to exploit those models, spanning from requirement specification to code development. Model-Driven Engineering (MDE) is a paradigm that leverages software models as primary artifacts to automate tasks. In this respect, modelers have started to investigate the interplay between traditional MDE practices and Large Language Models (LLMs) to push automation. Although powerful, LLMs exhibit limitations that undermine the quality of generated modeling artifacts, e.g., hallucination or incorrect formatting. Recording modeling operations relies on human-based activities to train modeling assistants, helping modelers in their daily tasks. Nevertheless, those techniques require a huge amount of training data that cannot be available due to several factors, e.g., security or privacy issues. Objective: In this paper, we propose an extension of a conceptual MDE framework, called MASTER-LLM, that combines different MDE tools and paradigms to support industrial and academic practitioners. Method: MASTER-LLM comprises a modeling environment that acts as the active context in which a dedicated component records modeling operations. Then, model completion is enabled by the modeling assistant trained on past operations. Different LLMs are used to generate a new dataset of modeling events to speed up recording and data collection. Results: To evaluate the feasibility of MASTER-LLM in practice, we experiment with two modeling environments, i.e., CAEX and HEPSYCODE, employed in industrial use cases within European projects. We investigate how the examined LLMs can generate realistic modeling operations in different domains. Conclusion: We show that synthetic traces can be effectively used when the application domain is less complex, while complex scenarios require human-based operations or a mixed approach according to data availability. However, generative AI models must be assessed using proper methodologies to avoid security issues in industrial domains. Vittoriano Muttillo, Claudio Di Sipio, Riccardo Rubei, Luca Berardinelli |
Inf. Softw. Technol. | 3 |
| 2025 | DeepMig: A transformer-based approach to support coupled library and code migrationsabstractWhile working on software projects, developers often replace third-party libraries (TPLs) with different ones offering similar functionalities. However, choosing a suitable TPL to migrate to is a complex task. As TPLs provide developers with Application Programming Interfaces (APIs) to allow for the invocation of their functionalities after adopting a new TPL, projects need to be migrated by the methods containing the affected API calls. Altogether, the coupled migration of TPLs and code is a strenuous process, requiring massive development effort. Most of the existing approaches either deal with library or API call migration but usually fail to solve both problems coherently simultaneously. This paper presents DeepMig, a novel approach to the coupled migration of TPLs and API calls. We aim to support developers in managing their projects, at the library and API level, allowing them to increase their productivity. DeepMig is based on a transformer architecture, accepts a set of libraries to predict a new set of libraries. Then, it looks for the changed API calls and recommends a migration plan for the affected methods. We evaluate DeepMig using datasets of Java projects collected from the Maven Central Repository, ensuring an assessment based on real-world dependency configurations. Our evaluation reveals promising outcomes: DeepMig recommends both libraries and code; by several projects, it retrieves a perfect match for the recommended items, obtaining an accuracy of 1.0. Moreover, being fed with proper training data, DeepMig provides comparable code migration steps of a static API migrator, a baseline for the code migration task. We conclude that DeepMig is capable of recommending both TPL and API migration, providing developers with a practical tool to migrate the entire project. • The migration of TPLs boils down to transforming of sequence of libraries. • Code migration is equal to the transition of sequences of API invocations. • Transformers can be used for migrating TPLs and APIs. • DeepMig recommends more relevant migration when there is enough data for learning. Juri Di Rocco, Phuong T. Nguyen 0001, Claudio Di Sipio, Riccardo Rubei, Davide Di Ruscio, Massimiliano Di Penta |
Inf. Softw. Technol. | 4 |
| 2025 | ModelXGlue: a benchmarking framework for ML tools in MDEabstractAbstract The integration of machine learning (ML) into model-driven engineering (MDE) holds the potential to enhance the efficiency of modelers and elevate the quality of modeling tools. However, a consensus is yet to be reached on which MDE tasks can derive substantial benefits from ML and how progress in these tasks should be measured. This paper introduces ModelXGlue , a dedicated benchmarking framework to empower researchers when constructing benchmarks for evaluating the application of ML to address MDE tasks. A benchmark is built by referencing datasets and ML models provided by other researchers, and by selecting an evaluation strategy and a set of metrics. ModelXGlue is designed with automation in mind and each component operates in an isolated execution environment (via Docker containers or Python environments), which allows the execution of approaches implemented with diverse technologies like Java, Python, R, etc. We used ModelXGlue to build reference benchmarks for three distinct MDE tasks: model classification, clustering, and feature name recommendation. To build the benchmarks we integrated existing third-party approaches in ModelXGlue . This shows that ModelXGlue is able to accommodate heterogeneous ML models, MDE tasks and different technological requirements. Moreover, we have obtained, for the first time, comparable results for these tasks. Altogether, it emerges that ModelXGlue is a valuable tool for advancing the understanding and evaluation of ML tools within the context of MDE. José Antonio Hernández López, Jesús Sánchez Cuadrado, Riccardo Rubei, Davide Di Ruscio |
Softw. Syst. Model. | 3 |
| 2025 | On the use of large language models in model-driven engineering
Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei |
Softw. Syst. Model. | 5 |
| 2025 | On the Energy Consumption of ATL TransformationsabstractAbstract Background Model transformations play a crucial role in Model‐Driven Engineering (MDE), with the ATLAS Transformation Language (ATL) being a powerful technology for developing model‐to‐model transformations. Methods This paper presents a comprehensive investigation into the energy consumption of ATL transformations, aiming to identify possible correlations among transformation rules, model size, and metamodel structural characteristics. We conducted experiments on 52 ATL transformations, analyzing power usage and extending our inquiry to understand the impact of mutations on both models and transformations. Results The experimental findings reveal relationships between the energy utilization of ATL transformations and the structural characteristics of metamodels. Furthermore, we establish a connection between energy consumption, model size, and the complexity of transformation processes. Conclusion The insights gained from this research lay the groundwork for devising future energy‐efficient strategies while developing model transformations. Riccardo Rubei, Juri Di Rocco, Davide Di Ruscio |
Softw. Pract. Exp. | 1 |
| 2024 | An Empirical Study on Code Coverage of Performance TestingabstractPerformance testing aims to ensure the operational efficiency of software systems. However, many factors influencing the efficacy and adoption of performance tests in practice are not yet fully understood. For instance, while code coverage is widely regarded as a key quality metric for evaluating the efficacy of functional testing suites, there is limited knowledge about the types and levels of coverage that performance tests specifically achieve. Another important factor, often perceived as a barrier to the broader adoption of performance tests yet remaining relatively unexplored, is their extended execution time. In this paper, we analyze the performance testing suites of 28 open-source systems to study (i) the magnitude of their code coverage, and (ii) their execution time. Our analysis shows that performance tests achieve significantly lower code coverage than functional tests, as expected, and it highlights a significant trade-off between coverage and execution time. Our results also suggest, in perspective, that automated test generation methods might not ensure affordable performance testing due to the associated time cost. This finding poses new challenges in the field of performance test generation. Muhammad Imran 0026, Vittorio Cortellessa, Davide Di Ruscio, Riccardo Rubei, Luca Traini |
EASE | 4 |
| 2024 | Automated categorization of pre-trained models in software engineering: A case study with a Hugging Face datasetabstractSoftware engineering (SE) activities have been revolutionized by the advent of pre-trained models (PTMs), defined as large machine learning (ML) models that can be fine-tuned to perform specific SE tasks. However, users with limited expertise may need help to select the appropriate model for their current task. To tackle the issue, the Hugging Face (HF) platform simplifies the use of PTMs by collecting, storing, and curating several models. Nevertheless, the platform currently lacks a comprehensive categorization of PTMs designed specifically for SE, i.e., the existing tags are more suited to generic ML categories. Claudio Di Sipio, Riccardo Rubei, Juri Di Rocco, Davide Di Ruscio, Phuong T. Nguyen 0001 |
EASE | 2 |
| 2024 | Towards Synthetic Trace Generation of Modeling Operations using In-Context Learning ApproachabstractProducing accurate software models is crucial in model-driven software engineering (MDE). However, modeling complex systems is an error-prone task that requires deep application domain knowledge. In the past decade, several automated techniques have been proposed to support academic and industrial practitioners by providing relevant modeling operations. Nevertheless, those techniques require a huge amount of training data that cannot be available due to several factors, e.g., privacy issues. The advent of large language models (LLMs) can support the generation of synthetic data although state-of-the-art approaches are not yet supporting the generation of modeling operations. To fill the gap, we propose a conceptual framework that combines modeling event logs, intelligent modeling assistants, and the generation of modeling operations using LLMs. In particular, the architecture comprises modeling components that help the designer specify the system, record its operation within a graphical modeling environment, and automatically recommend relevant operations. In addition, we generate a completely new dataset of modeling events by telling on the most prominent LLMs currently available. As a proof of concept, we instantiate the proposed framework using a set of existing modeling tools employed in industrial use cases within different European projects. To assess the proposed methodology, we first evaluate the capability of the examined LLMs to generate realistic modeling operations by relying on well-founded distance metrics. Then, we evaluate the recommended operations by considering real-world industrial modeling artifacts. Our findings demonstrate that LLMs can generate modeling events even though the overall accuracy is higher when considering human-based operations. In this respect, we see generative AI tools as an alternative when the modeling operations are not available to train traditional IMAs specifically conceived to support industrial practitioners. Vittoriano Muttillo, Claudio Di Sipio, Riccardo Rubei, Luca Berardinelli, MohammadHadi Dehghani |
ASE | 3 |
| 2024 | PlayMyData: a curated dataset of multi-platform video gamesabstractBeing predominant in digital entertainment for decades, video games have been recognized as valuable software artifacts by the software engineering (SE) community just recently. Such an acknowledgment has unveiled several research opportunities, spanning from empirical studies to the application of AI techniques for classification tasks. In this respect, several curated game datasets have been disclosed for research purposes even though the collected data are insufficient to support the application of advanced models or to enable interdisciplinary studies. Moreover, the majority of those are limited to PC games, thus excluding notorious gaming platforms, e.g., PlayStation, Xbox, and Nintendo. In this paper, we propose PlayMyData, a curated dataset composed of 99,864 multi-platform games gathered by the IGDB website. By exploiting a dedicated API, we collect relevant metadata for each game, e.g., description, genre, rating, gameplay video URLs, and screenshots. Furthermore, we enrich PlayMyData with the timing needed to complete each game by mining the HLTB website. To the best of our knowledge, this is the most comprehensive dataset in the domain that can be used to support different automated tasks in SE. More importantly, PlayMyData can be used to foster cross-domain investigations built on top of the provided multimedia data. Andrea D'Angelo, Claudio Di Sipio, Cristiano Politowski, Riccardo Rubei |
MSR | 4 |
| 2024 | GPTSniffer: A CodeBERT-based classifier to detect source code written by ChatGPTabstractSince its launch in November 2022, ChatGPT has gained popularity among users, especially programmers who use it to solve development issues. However, while offering a practical solution to programming problems, ChatGPT should be used primarily as a supporting tool (e.g., in software education) rather than as a replacement for humans. Thus, detecting automatically generated source code by ChatGPT is necessary, and tools for identifying AI-generated content need to be adapted to work effectively with code. This paper presents GPTSniffer– a novel approach to the detection of source code written by AI–built on top of CodeBERT. We conducted an empirical study to investigate the feasibility of automated identification of AI-generated code, and the factors that influence this ability. The results show that GPTSniffer can accurately classify whether code is human-written or AI-generated, outperforming two baselines, GPTZero and OpenAI Text Classifier. Also, the study shows how similar training data or a classification context with paired snippets helps boost the prediction. We conclude that GPTSniffer can be leveraged in different contexts, e.g., in software engineering education, where teachers use the tool to detect cheating and plagiarism, or in development, where AI-generated code may require peculiar quality assurance activities. Phuong T. Nguyen 0001, Juri Di Rocco, Claudio Di Sipio, Riccardo Rubei, Davide Di Ruscio, Massimiliano Di Penta |
J. Syst. Softw. | 4 |
| 2023 | Dealing with Popularity Bias in Recommender Systems for Third-party Libraries: How far Are We?abstractRecommender systems for software engineering (RSSEs) assist software engineers in dealing with a growing information overload when discerning alternative development solutions. While RSSEs are becoming more and more effective in suggesting handy recommendations, they tend to suffer from popularity bias, i.e., favoring items that are relevant mainly because several developers are using them. While this rewards artifacts that are likely more reliable and well-documented, it would also mean that missing artifacts are rarely used because they are very specific or more recent. This paper studies popularity bias in Third-Party Library (TPL) RSSEs. First, we investigate whether state-of-the-art research in RSSEs has already tackled the issue of popularity bias. Then, we quantitatively assess four existing TPL RSSEs, exploring their capability to deal with the recommendation of popular items. Finally, we propose a mechanism to defuse popularity bias in the recommendation list. The empirical study reveals that the issue of dealing with popularity in TPL RSSEs has not received adequate attention from the software engineering community. Among the surveyed work, only one starts investigating the issue, albeit getting a low prediction performance. Phuong T. Nguyen 0001, Riccardo Rubei, Juri Di Rocco, Claudio Di Sipio, Davide Di Ruscio, Massimiliano Di Penta |
MSR | 2 |
| 2023 | HybridRec: A recommender system for tagging GitHub repositoriesabstractAbstract Software repositories are increasingly essential to support the management of typical artifacts building up projects, including source code, documentation, and bug reports. GitHub is at the forefront of this kind of platforms, providing developer with a reservoir of code contained in more than 28M repositories. To help developers find the right artifacts, GitHub uses topics, which are short texts assigned to the stored artifacts. However, assigning inappropriate topics to a repository might hamper its popularity and reachability. In our previous work, we implemented MNBN and TopFilter to recommend GitHub topics. MNBN exploits a stochastic network to predict topics, while TopFilter relies on a syntactic-based function to recommend topics. In this paper, we extend our work by building HybridRec, a recommender system based on stochastic and collaborative-filtering techniques to generate more relevant topics. To deal with unbalanced datasets, we employ a Complement Naïve Bayesian Network (CNBN). Furthermore, we apply a preprocessing phase to clean and refine the input data before feeding the recommendation engine. An empirical evaluation demonstrates that HybridRec outperforms three state-of-the-art baselines, obtaining a better performance with respect to various metrics. We conclude that the conceived framework can be used to help developers increase their projects’ visibility. Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei |
Appl. Intell. | 5 |
| 2022 | Machine learning methods for model classification: a comparative studyabstractIn the quest to reuse modeling artifacts, academics and industry have proposed several model repositories over the last decade. Different storage and indexing techniques have been conceived to facilitate searching capabilities to help users find reusable artifacts that might fit the situation at hand. In this respect, machine learning (ML) techniques have been proposed to categorize and group large sets of modeling artifacts automatically. This paper reports the results of a comparative study of different ML classification techniques employed to automatically label models stored in model repositories. We have built a framework to systematically compare different ML models (feed-forward neural networks, graph neural networks, k-nearest neighbors, support version machines, etc.) with varying model encodings (TF-IDF, word embeddings, graphs and paths). We apply this framework to two datasets of about 5,000 Ecore and 5,000 UML models. We show that specific ML models and encodings perform better than others depending on the characteristics of the available datasets (e.g., the presence of duplicates) and on the goals to be achieved. José Antonio Hernández López, Riccardo Rubei, Jesús Sánchez Cuadrado, Davide Di Ruscio |
MoDELS | 2 |
| 2022 | Endowing third-party libraries recommender systems with explicit user feedback mechanismsabstractDuring their daily routine, developers often deal with a plethora of resources, attempting to search for relevant artifacts that can be added to the project under development. This kind of information overload may render developers overwhelmed, thus undermining their productivity and efficiency. Recommender systems are an effective means of easing such a burden, providing relevant items for the current programming contexts, e.g., third-party libraries (TPLs), API calls, or code snippets. By focusing on TPLs, there has been no work to allow for the integration of tailored feedback mechanisms with which users can conveniently accept or discard libraries. In this paper, we propose an approach to handle explicit user feedback, including positive, negative, and additive. Thus, further than accepting or discarding the recommended TPLs, users can also endorse libraries that, in their opinion, are relevant for the current context, even though they are not included in the provided recommendations. As a proof of concept, we demonstrate how user feedback generated by the proposed mechanism can change the outcome of a real TPLs recommender system. The results show that our proposed approach helps the considered system retrieve relevant items, under different configurations. Riccardo Rubei, Claudio Di Sipio, Juri Di Rocco, Davide Di Ruscio, Phuong T. Nguyen 0001 |
SANER | 1 |
| 2022 | Providing upgrade plans for third-party libraries: a recommender system using migration graphs
Riccardo Rubei, Davide Di Ruscio, Claudio Di Sipio, Juri Di Rocco, Phuong T. Nguyen 0001 |
Appl. Intell. | 1 |
| 2022 | DeepLib: Machine translation techniques to recommend upgrades for third-party libraries
Phuong T. Nguyen 0001, Juri Di Rocco, Riccardo Rubei, Claudio Di Sipio, Davide Di Ruscio |
Expert Syst. Appl. | 3 |
| 2021 | Development of recommendation systems for software engineering: the CROSSMINER experienceabstractAbstract To perform their daily tasks, developers intensively make use of existing resources by consulting open source software (OSS) repositories. Such platforms contain rich data sources, e.g., code snippets, documentations, and user discussions, that can be useful for supporting development activities. Over the last decades, several techniques and tools have been promoted to provide developers with innovative features, aiming to bring in improvements in terms of development effort, cost savings, and productivity. In the context of the EU H2020 CROSSMINER project, a set of recommendation systems has been conceived to assist software programmers in different phases of the development process. The systems provide developers with various artifacts, such as third-party libraries, documentation about how to use the APIs being adopted, or relevant API function calls. To develop such recommendations, various technical choices have been made to overcome issues related to several aspects including the lack of baselines, limited data availability, decisions about the performance measures, and evaluation approaches. This paper is an experience report to present the knowledge pertinent to the set of recommendation systems developed through the CROSSMINER project. We explain in detail the challenges we had to deal with, together with the related lessons learned when developing and evaluating these systems. Our aim is to provide the research community with concrete takeaway messages that are expected to be useful for those who want to develop or customize their own recommendation systems. The reported experiences can facilitate interesting discussions and research work, which in the end contribute to the advancement of recommendation systems applied to solve different issues in Software Engineering. Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei |
Empir. Softw. Eng. | 5 |
| 2020 | A Multinomial Naïve Bayesian (MNB) Network to Automatically Recommend Topics for GitHub RepositoriesabstractGitHub has become a precious service for storing and managing software source code. Over the last year, 10M new developers have joined the GitHub community, contributing to more than 44M repositories. In order to help developers increase the reachability of their repositories, in 2017 GitHub introduced the possibility to classify them by means of topics. However, assigning wrong topics to a given repository can compromise the possibility of helping other developers approach it, and thus preventing them from contributing to its development. Claudio Di Sipio, Riccardo Rubei, Davide Di Ruscio, Phuong T. Nguyen 0001 |
EASE | 2 |
| 2020 | TopFilter: An Approach to Recommend Relevant GitHub TopicsabstractBackground: In the context of software development, GitHub has been at the forefront of platforms to store, analyze and maintain a large number of software repositories. Topics have been introduced by GitHub as an effective method to annotate stored repositories. However, labeling GitHub repositories should be carefully conducted to avoid adverse effects on project popularity and reachability. Aims: We present TopFilter, a novel approach to assist open source software developers in selecting suitable topics for GitHub repositories being created. Method: We built a project-topic matrix and applied a syntactic-based similarity function to recommend missing topics by representing repositories and related topics in a graph. The ten-fold cross-validation methodology has been used to assess the performance of TopFilter by considering different metrics, i.e., success rate, precision, recall, and catalog coverage. Result: The results show that TopFilter recommends good topics depending on different factors, i.e., collaborative filtering settings, considered datasets, and pre-processing activities. Moreover, TopFilter can be combined with a state-of-the-art topic recommender system (i.e., MNB network) to improve the overall prediction performance. Conclusion: Our results confirm that collaborative filtering techniques can successfully be used to provide relevant topics for GitHub repositories. Moreover, TopFilter can gain a significant boost in prediction performances by employing the outcomes obtained by the MNB network as its initial set of topics. Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei |
ESEM | 5 |
| 2020 | PostFinder: Mining Stack Overflow posts to support software developers
Riccardo Rubei, Claudio Di Sipio, Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio |
Inf. Softw. Technol. | 1 |
| 2020 | An automated approach to assess the similarity of GitHub repositories
Phuong T. Nguyen 0001, Juri Di Rocco, Riccardo Rubei, Davide Di Ruscio |
Softw. Qual. J. | 3 |
| 2018 | CrossSim: Exploiting Mutual Relationships to Detect Similar OSS ProjectsabstractSoftware development is a knowledge-intensive activity, which requires mastering several languages, frameworks, technology trends (among other aspects) under the pressure of ever-increasing arrays of external libraries and resources. Recommender systems are gaining high relevance in software engineering since they aim at providing developers with real-time recommendations, which can reduce the time spent on discovering and understanding reusable artifacts from software repositories, and thus inducing productivity and quality gains. In this paper, we focus on the problem of mining open source software repositories to identify similar projects, which can be evaluated and eventually reused by developers. To this end, CrossSim is proposed as a novel approach to model open source software projects and related artifacts and to compute similarities among them. An evaluation on a dataset containing 580 GitHub projects shows that CrossSim outperforms an existing technique, which has been proven to have a good performance in detecting similar GitHub repositories. Phuong T. Nguyen 0001, Juri Di Rocco, Riccardo Rubei, Davide Di Ruscio |
SEAA | 3 |