Juri Di Rocco

dblp:145/3999 · DBLP profile ↗
← Back
51ranked-venue papers
12as first author
33since 2021 · last 2026
0000-0002-7909-3902ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 46 · 11 first-author · 28 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Automated summarization of software documents: an LLM-based multi-agent approach
Duc S. H. Nguyen, Minh T. Nguyen, Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio
Autom. Softw. Eng.4
2025 Generate with CodeXHug: A Dataset to Enhance Model Cards with Code Usage Patterns
Stefano Palombo, Claudio Di Sipio, Juri Di Rocco, Davide Di Ruscio
SEAA (3)3
2025 Binary and multi-class classification of Self-Admitted Technical Debt: How far can we go?
abstract
Context: Aiming for a trade-off between short-term efficiency and long-term stability, software teams resort to sub-optimal solutions, neglecting the best software development practices. Such solutions may induce technical debt (TD), triggering maintenance issues. To facilitate future fixing, developers mark code with any issues using textual comments, resulting in Self-Admitted Technical Debt (SATD). Detecting SATD in source code is crucial since it helps programmers locate potentially erroneous snippets, allowing for suitable interventions, and improving code quality. There are two main types of SATD detection, i.e., binary classification and multi-class classification , grouping TD comments into SATD/Non-SATD categories, and multiple categories, respectively. Objective: We attempt to understand to which extent state-of-the-art research has addressed the issue of detecting SATD, both binary and multi-class classification. Based on this investigation, we also propose a practical approach for the detection of SATD using Large Language Models (LLMs). Methods: First, we conducted a literature review to understand to which extent the two types of classification have been tackled by existing research. Second, we developed SALA , a dual-purpose tool on top of Natural Language Processing (NLP) techniques and neural networks to deal with both types of classification. An empirical evaluation has been performed to compare SALA with state-of-the-art baselines. Results: The literature review reveals that while binary classification has been well studied, multi-class classification has not received adequate attention. The empirical evaluation shows that SALA obtains a promising performance, and outperforms the baselines with respect to various quality metrics. Conclusion: We conclude that more effort needs to be spent to tackle multi-class classification of SATD. To this end, LLMs hold the potential, albeit with more rigorous investigation on possible fine-tuning and prompt engineering strategies.
Francesca Arcelli Fontana, Juri Di Rocco, Davide Di Ruscio, Amleto Di Salle, Phuong T. Nguyen 0001
Inf. Softw. Technol.2
2025 DeepMig: A transformer-based approach to support coupled library and code migrations
abstract
While working on software projects, developers often replace third-party libraries (TPLs) with different ones offering similar functionalities. However, choosing a suitable TPL to migrate to is a complex task. As TPLs provide developers with Application Programming Interfaces (APIs) to allow for the invocation of their functionalities after adopting a new TPL, projects need to be migrated by the methods containing the affected API calls. Altogether, the coupled migration of TPLs and code is a strenuous process, requiring massive development effort. Most of the existing approaches either deal with library or API call migration but usually fail to solve both problems coherently simultaneously. This paper presents DeepMig, a novel approach to the coupled migration of TPLs and API calls. We aim to support developers in managing their projects, at the library and API level, allowing them to increase their productivity. DeepMig is based on a transformer architecture, accepts a set of libraries to predict a new set of libraries. Then, it looks for the changed API calls and recommends a migration plan for the affected methods. We evaluate DeepMig using datasets of Java projects collected from the Maven Central Repository, ensuring an assessment based on real-world dependency configurations. Our evaluation reveals promising outcomes: DeepMig recommends both libraries and code; by several projects, it retrieves a perfect match for the recommended items, obtaining an accuracy of 1.0. Moreover, being fed with proper training data, DeepMig provides comparable code migration steps of a static API migrator, a baseline for the code migration task. We conclude that DeepMig is capable of recommending both TPL and API migration, providing developers with a practical tool to migrate the entire project. • The migration of TPLs boils down to transforming of sequence of libraries. • Code migration is equal to the transition of sequences of API invocations. • Transformers can be used for migrating TPLs and APIs. • DeepMig recommends more relevant migration when there is enough data for learning.
Juri Di Rocco, Phuong T. Nguyen 0001, Claudio Di Sipio, Riccardo Rubei, Davide Di Ruscio, Massimiliano Di Penta
Inf. Softw. Technol.1
2025 On the use of large language models in model-driven engineering
Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei
Softw. Syst. Model.1
2025 A framework for evaluating tool support for co-evolution of modeling languages, tools and models
abstract
Abstract We present a framework for evaluating language workbenches’ capabilities for co-evolution of graphical modeling languages, modeling tools and models. As with programming, maintenance tasks such as language refinement and enhancement typically account for more work than the initial development phase. Modeling languages have the added challenge of keeping tools and existing models in step with the evolving language. As domain-specific modeling languages and tools have started to be used widely, thanks to reports of significant productivity improvements, some language workbench users have indeed reported problems with co-evolution of tools and models. Our tool-agnostic evaluation framework aims to cover changes across the whole language definition: the abstract syntax, concrete syntax, and constraints. Change impact is assessed for knock-on effects within the language definition, the modeling tools, semantics via generators, and existing models. We demonstrate the viability of the framework by evaluating MetaEdit+, EMF/Sirius and Jjodel, providing a detailed evaluation process for others to repeat with their tools. The results of the evaluation show differences among the tools: from editors not opening correctly or at all, through highlighting of items requiring manual intervention, to fully automatic updates of languages, models and editors. We call for industry to evaluate their tool choices with the framework, tool developers to extend their tool support for co-evolution, and researchers to refine the evaluation framework and evaluations presented.
Juha-Pekka Tolvanen, Steven Kelly 0001, Juri Di Rocco, Alfonso Pierantonio, Giordano Tinella
Softw. Syst. Model.3
2025 On the Energy Consumption of ATL Transformations
abstract
Abstract Background Model transformations play a crucial role in Model‐Driven Engineering (MDE), with the ATLAS Transformation Language (ATL) being a powerful technology for developing model‐to‐model transformations. Methods This paper presents a comprehensive investigation into the energy consumption of ATL transformations, aiming to identify possible correlations among transformation rules, model size, and metamodel structural characteristics. We conducted experiments on 52 ATL transformations, analyzing power usage and extending our inquiry to understand the impact of mutations on both models and transformations. Results The experimental findings reveal relationships between the energy utilization of ATL transformations and the structural characteristics of metamodels. Furthermore, we establish a connection between energy consumption, model size, and the complexity of transformation processes. Conclusion The insights gained from this research lay the groundwork for devising future energy‐efficient strategies while developing model transformations.
Riccardo Rubei, Juri Di Rocco, Davide Di Ruscio
Softw. Pract. Exp.2
2024 Automated categorization of pre-trained models in software engineering: A case study with a Hugging Face dataset
abstract
Software engineering (SE) activities have been revolutionized by the advent of pre-trained models (PTMs), defined as large machine learning (ML) models that can be fine-tuned to perform specific SE tasks. However, users with limited expertise may need help to select the appropriate model for their current task. To tackle the issue, the Hugging Face (HF) platform simplifies the use of PTMs by collecting, storing, and curating several models. Nevertheless, the platform currently lacks a comprehensive categorization of PTMs designed specifically for SE, i.e., the existing tags are more suited to generic ML categories.
Claudio Di Sipio, Riccardo Rubei, Juri Di Rocco, Davide Di Ruscio, Phuong T. Nguyen 0001
EASE3
2024 Automatic Categorization of GitHub Actions with Transformers and Few-shot Learning
abstract
In the GitHub ecosystem, workflows are used as an effective means to automate development tasks and to set up a Continuous Integration and Delivery (CI/CD pipeline). GitHub Actions (GHA) has been conceived to provide developers with a practical tool to create and maintain workflows, avoiding “reinventing the wheel” and cluttering the workflow with shell commands. Properly leveraging the power of GitHub Actions can facilitate the development processes, enhance collaboration, and significantly impact project outcomes. To expose actions to search engines, GitHub allows developers to assign them to one or more categories manually. These are used as an effective means to group actions sharing similar functionality. Nevertheless, while providing a practical way to execute workflows, many actions have unclear purposes, and sometimes they are not categorized. In this work, we bridge such a gap by conceptualizing Gavel, a practical solution to increasing the visibility of actions in GitHub. By leveraging the content of README.MD files for each action, we use Transformer to assign suitable categories to the action. We conducted an empirical investigation and compared Gavel with a state-of-the-art baseline. The results show that our approach can assign categories to GitHub actions effectively, thus outperforming the baseline.
Phuong T. Nguyen 0001, Juri Di Rocco, Claudio Di Sipio, Mudita Shakya, Davide Di Ruscio, Massimiliano Di Penta
ESEM2
2024 Exploring user privacy awareness on GitHub: an empirical study
abstract
Abstract GitHub provides developers with a practical way to distribute source code and collaboratively work on common projects. To enhance account security and privacy, GitHub allows its users to manage access permissions, review audit logs, and enable two-factor authentication. However, despite the endless effort, the platform still faces various issues related to the privacy of its users. This paper presents an empirical study delving into the GitHub ecosystem. Our focus is on investigating the utilization of privacy settings on the platform and identifying various types of sensitive information disclosed by users. Leveraging a dataset comprising 6,132 developers, we report and analyze their activities by means of comments on pull requests. Our findings indicate an active engagement by users with the available privacy settings on GitHub. Notably, we observe the disclosure of different forms of private information within pull request comments. This observation has prompted our exploration into sensitivity detection using a large language model and BERT, to pave the way for a personalized privacy assistant. Our work provides insights into the utilization of existing privacy protection tools, such as privacy settings, along with their inherent limitations. Essentially, we aim to advance research in this field by providing both the motivation for creating such privacy protection tools and a proposed methodology for personalizing them.
Costanza Alfieri, Juri Di Rocco, Paola Inverardi, Phuong T. Nguyen 0001
Empir. Softw. Eng.2
2024 GPTSniffer: A CodeBERT-based classifier to detect source code written by ChatGPT
abstract
Since its launch in November 2022, ChatGPT has gained popularity among users, especially programmers who use it to solve development issues. However, while offering a practical solution to programming problems, ChatGPT should be used primarily as a supporting tool (e.g., in software education) rather than as a replacement for humans. Thus, detecting automatically generated source code by ChatGPT is necessary, and tools for identifying AI-generated content need to be adapted to work effectively with code. This paper presents GPTSniffer– a novel approach to the detection of source code written by AI–built on top of CodeBERT. We conducted an empirical study to investigate the feasibility of automated identification of AI-generated code, and the factors that influence this ability. The results show that GPTSniffer can accurately classify whether code is human-written or AI-generated, outperforming two baselines, GPTZero and OpenAI Text Classifier. Also, the study shows how similar training data or a classification context with paired snippets helps boost the prediction. We conclude that GPTSniffer can be leveraged in different contexts, e.g., in software engineering education, where teachers use the tool to detect cheating and plagiarism, or in development, where AI-generated code may require peculiar quality assurance activities.
Phuong T. Nguyen 0001, Juri Di Rocco, Claudio Di Sipio, Riccardo Rubei, Davide Di Ruscio, Massimiliano Di Penta
J. Syst. Softw.2
2024 TyphonML: Tool support for hybrid polystores
Francesco Basciani, Juri Di Rocco, Ludovico Iovino, Alfonso Pierantonio
Sci. Comput. Program.2
2024 Advanced discovery mechanisms in model repositories
abstract
Summary As model‐driven engineering gains traction and poses as the new paradigm for software engineering, it raises a need for efficient approaches and tools to manage, discover, and retrieve relevant modeling artifacts. Hence, industry and academia are conceiving effective ways to store, search, and retrieve heterogeneous model artifacts that employ advanced discovery mechanisms. This paper presents MDEForge‐Search, a novel approach to discovering heterogeneous model artifacts over MDEForge, a distributed cloud‐based model repository. We designed advanced discovery mechanisms that retrieve heterogeneous artifacts within their context (megamodel) and reuse them across model management services. In addition, a domain‐specific approach has been proposed to formulate queries in terms of keywords, search tags, conditional operators, quality model assessment services and a transformation chain discoverer. Finally, the applicability of our approach was assessed in a recommender system modeling framework, which, thanks to the operated integration, can rely on the availability of more than 5000 model artifacts currently persisted in our cloud‐based model repository.
Arsene Indamutsa, Juri Di Rocco, Lissette Almonte, Davide Di Ruscio, Alfonso Pierantonio
Softw. Pract. Exp.2
2023 Too long; didn't read: Automatic summarization of GitHub README.MD with Transformers
abstract
The ability to allow developers to share their source code and collaborate on software projects has made GitHub a widely used open source platform. Each repository in GitHub is generally equipped with a README.MD file to exhibit an overview of the main functionalities. Nevertheless, while offering useful information, README.MD is usually lengthy, requiring time and effort to read and comprehend. Thus, besides README.MD, GitHub also allows its users to add a short description called “About,” giving a brief but informative summary about the repository. This enables visitors to quickly grasp the main content and decide whether to continue reading. Unfortunately, due to various reasons–not excluding laziness–oftentimes this field is left blank by developers.
Thu Thu Ha Doan, Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio
EASE3
2023 Dealing with Popularity Bias in Recommender Systems for Third-party Libraries: How far Are We?
abstract
Recommender systems for software engineering (RSSEs) assist software engineers in dealing with a growing information overload when discerning alternative development solutions. While RSSEs are becoming more and more effective in suggesting handy recommendations, they tend to suffer from popularity bias, i.e., favoring items that are relevant mainly because several developers are using them. While this rewards artifacts that are likely more reliable and well-documented, it would also mean that missing artifacts are rarely used because they are very specific or more recent. This paper studies popularity bias in Third-Party Library (TPL) RSSEs. First, we investigate whether state-of-the-art research in RSSEs has already tackled the issue of popularity bias. Then, we quantitatively assess four existing TPL RSSEs, exploring their capability to deal with the recommendation of popular items. Finally, we propose a mechanism to defuse popularity bias in the recommendation list. The empirical study reveals that the issue of dealing with popularity in TPL RSSEs has not received adequate attention from the software engineering community. Among the surveyed work, only one starts investigating the issue, albeit getting a low prediction performance.
Phuong T. Nguyen 0001, Riccardo Rubei, Juri Di Rocco, Claudio Di Sipio, Davide Di Ruscio, Massimiliano Di Penta
MSR3
2023 HybridRec: A recommender system for tagging GitHub repositories
abstract
Abstract Software repositories are increasingly essential to support the management of typical artifacts building up projects, including source code, documentation, and bug reports. GitHub is at the forefront of this kind of platforms, providing developer with a reservoir of code contained in more than 28M repositories. To help developers find the right artifacts, GitHub uses topics, which are short texts assigned to the stored artifacts. However, assigning inappropriate topics to a repository might hamper its popularity and reachability. In our previous work, we implemented MNBN and TopFilter to recommend GitHub topics. MNBN exploits a stochastic network to predict topics, while TopFilter relies on a syntactic-based function to recommend topics. In this paper, we extend our work by building HybridRec, a recommender system based on stochastic and collaborative-filtering techniques to generate more relevant topics. To deal with unbalanced datasets, we employ a Complement Naïve Bayesian Network (CNBN). Furthermore, we apply a preprocessing phase to clean and refine the input data before feeding the recommendation engine. An empirical evaluation demonstrates that HybridRec outperforms three state-of-the-art baselines, obtaining a better performance with respect to various metrics. We conclude that the conceived framework can be used to help developers increase their projects’ visibility.
Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei
Appl. Intell.1
2023 Fitting missing API puzzles with machine translation techniques
Phuong T. Nguyen 0001, Claudio Di Sipio, Juri Di Rocco, Davide Di Ruscio, Massimiliano Di Penta
Expert Syst. Appl.3
2023 MemoRec: a recommender system for assisting modelers in specifying metamodels
abstract
Abstract Model-driven engineering has been widely applied in software development, aiming to facilitate the coordination among various stakeholders. Such a methodology allows for a more efficient and effective development process. Nevertheless, modeling is a strenuous activity that requires proper knowledge of components, attributes, and logic to reach the level of abstraction required by the application domain. In particular, metamodels play an important role in several paradigms, and specifying wrong entities or attributes in metamodels can negatively impact on the quality of the produced artifacts as well as other elements of the whole process. During the metamodeling phase, modelers can benefit from assistance to avoid mistakes, e.g., getting recommendations like metaclasses and structural features relevant to the metamodel being defined. However, suitable machinery is needed to mine data from repositories of existing modeling artifacts and compute recommendations. In this work, we propose MemoRec, a novel approach that makes use of a collaborative filtering strategy to recommend valuable entities related to the metamodel under construction. Our approach can provide suggestions related to both metaclasses and structured features that should be added in the metamodel under definition. We assess the quality of the work with respect to different metrics, i.e., success rate, precision, and recall. The results demonstrate that MemoRec is capable of suggesting relevant items given a partial metamodel and supporting modelers in their task.
Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Alfonso Pierantonio
Softw. Syst. Model.1
2023 MORGAN: a modeling recommender system based on graph kernel
abstract
Abstract Model-driven engineering (MDE) is an effective means of synchronizing among stakeholders, thereby being a crucial part of the software development life cycle. In recent years, MDE has been on the rise, triggering the need for automatic modeling assistants to support metamodelers during their daily activities. Among others, it is crucial to enable model designers to choose suitable components while working on new (meta)models. In our previous work, we proposed MORGAN, a graph kernel-based recommender system to assist developers in completing models and metamodels. To provide input for the recommendation engine, we convert training data into a graph-based format, making use of various natural language processing (NLP) techniques. The extracted graphs are then fed as input for a recommendation engine based on graph kernel similarity, which performs predictions to provide modelers with relevant recommendations to complete the partially specified (meta)models. In this paper, we extend the proposed tool in different dimensions, resulting in a more advanced recommender system. Firstly, we equip it with the ability to support recommendations for JSON schema that provides a model representation of data handling operations. Secondly, we introduce additional preprocessing steps and a kernel similarity function based on item frequency, aiming to enhance the capabilities, providing more precise recommendations. Thirdly, we study the proposed enhancements, conducting a well-structured evaluation by considering three real-world datasets. Although the increasing size of the training data negatively affects the computation time, the experimental results demonstrate that the newly introduced mechanisms allow MORGAN to improve its recommendations compared to its preceding version.
Claudio Di Sipio, Juri Di Rocco, Davide Di Ruscio, Phuong T. Nguyen 0001
Softw. Syst. Model.2
2023 GitRanking: A ranking of GitHub topics for software classification using active sampling
abstract
Abstract Context GitHub is the world's most prominent host of source code, with more than 327M repositories. However, most of these repositories are not labelled or inadequately, making it harder for users to find relevant projects. Various proposals for software application domain classification over the past years have been proposed. However, these several of those approaches suffer from multiple issues, called antipatterns of software classification, that reduce their usability. Objective In this paper, we propose a new taxonomy in the GitHub ecosystem, called GitRanking, starting from a well‐structured data set, composed of curated repositories annotated with topics. The main objective is to create a baseline methodology for software classification that is expandable, hierarchical, grounded in a knowledge base, and free of antipatterns. Method We collected 121K topics from GitHub and used GitRanking to create a taxonomy of 301 ranked application domains. GitRanking (1) uses active sampling to ensure a minimal number of annotations to create the ranking; and (2) links each topic to Wikidata, reducing ambiguities and improving the reusability of the taxonomy. Furthermore, we adopt the conceived taxonomy in a classification task by considering a state‐of‐the‐art classifier. Results Our results show that GitRanking can effectively rank terms in a hierarchy according to how general or specific their meaning is. Furthermore, we show that GitRanking is a dynamically extensible method: it can currently accept further terms to be ranked, and with a minimum number of annotations (). Concerning the classification task, we show that the model achieves an F1‐score of 34%, with a precision of 54%. Conclusion This paper is the first collective attempt at building a ground‐up taxonomy of software domains. Our vision is that our taxonomy, and its extensibility, can be used to better and more precisely label software projects.
Cezar Sas, Andrea Capiluppi, Claudio Di Sipio, Juri Di Rocco, Davide Di Ruscio
Softw. Pract. Exp.4
2022 Finding with NEMO: a recommender system to forecast the next modeling operations
abstract
Nowadays, while modeling environments provide users with facilities to specify different kinds of artifacts, e.g., metamodels, models, and transformations, the possibility of learning from previous modeling experiences and being assisted during modeling tasks remains largely unexplored. In this paper, we propose NEMO, a recommender system based on an Encoder-Decoder neural network to assist modelers in performing model editing operations. NEMO learns from past modeling activities and performs predictions employing a deep learning technique. Such an algorithm has been successfully applied in machine translation to convert a text from a language to another foreign language and vice versa. An empirical evaluation on a dataset of BPMN change-based persistent model demonstrates that the technique permits learning from existing operations and effectively predicting the next editing operations with considerably high prediction accuracy. In particular, NEMO gets 0.977 as precision/recall and 0.992 as success rate score by the best performance.
Juri Di Rocco, Claudio Di Sipio, Phuong T. Nguyen 0001, Davide Di Ruscio, Alfonso Pierantonio
MoDELS1
2022 Endowing third-party libraries recommender systems with explicit user feedback mechanisms
abstract
During their daily routine, developers often deal with a plethora of resources, attempting to search for relevant artifacts that can be added to the project under development. This kind of information overload may render developers overwhelmed, thus undermining their productivity and efficiency. Recommender systems are an effective means of easing such a burden, providing relevant items for the current programming contexts, e.g., third-party libraries (TPLs), API calls, or code snippets. By focusing on TPLs, there has been no work to allow for the integration of tailored feedback mechanisms with which users can conveniently accept or discard libraries. In this paper, we propose an approach to handle explicit user feedback, including positive, negative, and additive. Thus, further than accepting or discarding the recommended TPLs, users can also endorse libraries that, in their opinion, are relevant for the current context, even though they are not included in the provided recommendations. As a proof of concept, we demonstrate how user feedback generated by the proposed mechanism can change the outcome of a real TPLs recommender system. The results show that our proposed approach helps the considered system retrieve relevant items, under different configurations.
Riccardo Rubei, Claudio Di Sipio, Juri Di Rocco, Davide Di Ruscio, Phuong T. Nguyen 0001
SANER3
2022 Providing upgrade plans for third-party libraries: a recommender system using migration graphs
Riccardo Rubei, Davide Di Ruscio, Claudio Di Sipio, Juri Di Rocco, Phuong T. Nguyen 0001
Appl. Intell.4
2022 DeepLib: Machine translation techniques to recommend upgrades for third-party libraries
Phuong T. Nguyen 0001, Juri Di Rocco, Riccardo Rubei, Claudio Di Sipio, Davide Di Ruscio
Expert Syst. Appl.2
2022 Recommending API Function Calls and Code Snippets to Support Software Development
abstract
Software development activity has reached a high degree of complexity, guided by the heterogeneity of the components, data sources, and tasks. The proliferation of open-source software (OSS) repositories has stressed the need to reuse available software artifacts efficiently. To this aim, it is necessary to explore approaches to mine data from software repositories and leverage it to produce helpful recommendations. We designed and implemented FOCUS as a novel approach to provide developers with API calls and source code while they are programming. The system works on the basis of a context-aware collaborative filtering technique to extract API usages from OSS projects. In this work, we show the suitability of FOCUS for Android programming by evaluating it on a dataset of 2,600 mobile apps. The empirical evaluation results show that our approach outperforms two state-of-the-art API recommenders, UP-Miner and PAM, in terms of prediction accuracy. We also point out that there is no significant relationship between the categories for apps defined in Google Play and their API usages. Finally, we show that participants of a user study positively perceive the API and source code recommended by FOCUS as relevant to the current development context.
Phuong T. Nguyen 0001, Juri Di Rocco, Claudio Di Sipio, Davide Di Ruscio, Massimiliano Di Penta
IEEE Trans. Software Eng.2
2021 Adversarial Machine Learning: On the Resilience of Third-party Library Recommender Systems
abstract
In recent years, we have witnessed a dramatic increase in the application of Machine Learning algorithms in several domains, including the development of recommender systems for software engineering (RSSE). While researchers focused on the underpinning ML techniques to improve recommendation accuracy, little attention has been paid to make such systems robust and resilient to malicious data. By manipulating the algorithms’ training set, i.e., large open-source software (OSS) repositories, it would be possible to make recommender systems vulnerable to adversarial attacks. This paper presents an initial investigation of adversarial machine learning and its possible implications on RSSE. As a proof-of-concept, we show the extent to which the presence of manipulated data can have a negative impact on the outcomes of two state-of-the-art recommender systems which suggest third-party libraries to developers. Our work aims at raising awareness of adversarial techniques and their effects on the Software Engineering community. We also propose equipping recommender systems with the capability to learn to dodge adversarial activities.
Phuong T. Nguyen 0001, Davide Di Ruscio, Juri Di Rocco, Claudio Di Sipio, Massimiliano Di Penta
EASE3
2021 Adversarial Attacks to API Recommender Systems: Time to Wake Up and Smell the Coffeeƒ
abstract
Recommender systems in software engineering provide developers with a wide range of valuable items to help them complete their tasks. Among others, API recommender systems have gained momentum in recent years as they became more successful at suggesting API calls or code snippets. While these systems have proven to be effective in terms of prediction accuracy, there has been less attention for what concerns such recommenders’ resilience against adversarial attempts. In fact, by crafting the recommenders’ learning material, e.g., data from large open-source software (OSS) repositories, hostile users may succeed in injecting malicious data, putting at risk the software clients adopting API recommender systems. In this paper, we present an empirical investigation of adversarial machine learning techniques and their possible influence on recommender systems. The evaluation performed on three state-of-the-art API recommender systems reveals a worrying outcome: all of them are not immune to malicious data. The obtained result triggers the need for effective countermeasures to protect recommender systems against hostile attacks disguised in training data.
Phuong T. Nguyen 0001, Claudio Di Sipio, Juri Di Rocco, Massimiliano Di Penta, Davide Di Ruscio
ASE3
2021 A GNN-based Recommender System to Assist the Specification of Metamodels and Models
abstract
Nowadays, while modeling environments provide users with facilities to specify different kinds of artifacts, e.g., metamodels, models, and transformations, the possibility of learning from previous modeling experiences and being assisted during modeling tasks remains largely unexplored. In this paper, we propose MORGAN, a recommender system based on a graph neural network (GNN) to assist modelers in performing the specification of metamodels and models. The (meta)model being specified, and the training data are encoded in a graph-based format by exploiting natural language processing (NLP) techniques. Afterward, a graph kernel function uses the extracted graphs to provide modelers with relevant recommendations to complete the partially specified (meta)models. We evaluated MORGAN on real-world datasets using various quality metrics, i.e., precision, recall, and F-measure. The experimental results are encouraging and demonstrate the feasibility of our tool to support modelers while specifying metamodels and models.
Juri Di Rocco, Claudio Di Sipio, Davide Di Ruscio, Phuong T. Nguyen 0001
MoDELS1
2021 A Low-Code Tool Supporting the Development of Recommender Systems
abstract
The design of recommender systems (RSs) to support software development encompasses the fulfillment of different steps, including data preprocessing, choice of the most appropriate algorithms, item delivery. Though RSs can alleviate the curse of information overload, existing approaches resemble black-box systems, in which the end-user is not expected to fine-tune or personalize the overall process.
Claudio Di Sipio, Juri Di Rocco, Davide Di Ruscio, Phuong T. Nguyen 0001
RecSys2
2021 Development of recommendation systems for software engineering: the CROSSMINER experience
abstract
Abstract To perform their daily tasks, developers intensively make use of existing resources by consulting open source software (OSS) repositories. Such platforms contain rich data sources, e.g., code snippets, documentations, and user discussions, that can be useful for supporting development activities. Over the last decades, several techniques and tools have been promoted to provide developers with innovative features, aiming to bring in improvements in terms of development effort, cost savings, and productivity. In the context of the EU H2020 CROSSMINER project, a set of recommendation systems has been conceived to assist software programmers in different phases of the development process. The systems provide developers with various artifacts, such as third-party libraries, documentation about how to use the APIs being adopted, or relevant API function calls. To develop such recommendations, various technical choices have been made to overcome issues related to several aspects including the lack of baselines, limited data availability, decisions about the performance measures, and evaluation approaches. This paper is an experience report to present the knowledge pertinent to the set of recommendation systems developed through the CROSSMINER project. We explain in detail the challenges we had to deal with, together with the related lessons learned when developing and evaluating these systems. Our aim is to provide the research community with concrete takeaway messages that are expected to be useful for those who want to develop or customize their own recommendation systems. The reported experiences can facilitate interesting discussions and research work, which in the end contribute to the advancement of recommendation systems applied to solve different issues in Software Engineering.
Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei
Empir. Softw. Eng.1
2021 Convolutional neural networks for enhanced classification mechanisms of metamodels
Phuong T. Nguyen 0001, Davide Di Ruscio, Alfonso Pierantonio, Juri Di Rocco, Ludovico Iovino
J. Syst. Softw.4
2021 Evaluation of a machine learning classifier for metamodels
abstract
Abstract Modeling is a ubiquitous activity in the process of software development. In recent years, such an activity has reached a high degree of intricacy, guided by the heterogeneity of the components, data sources, and tasks. The democratized use of models has led to the necessity for suitable machinery for mining modeling repositories. Among others, the classification of metamodels into independent categories facilitates personalized searches by boosting the visibility of metamodels. Nevertheless, the manual classification of metamodels is not only a tedious but also an error-prone task. According to our observation, misclassification is the norm which leads to a reduction in reachability as well as reusability of metamodels. Handling such complexity requires suitable tooling to leverage raw data into practical knowledge that can help modelers with their daily tasks. In our previous work, we proposed AURORA as a machine learning classifier for metamodel repositories. In this paper, we present a thorough evaluation of the system by taking into consideration different settings as well as evaluation metrics. More importantly, we improve the original AURORA tool by changing its internal design. Experimental results demonstrate that the proposed amendment is beneficial to the classification of metamodels. We also compared our approach with two baseline algorithms, namely gradient boosted decision tree and support vector machines. Eventually, we see that AURORA outperforms the baselines with respect to various quality metrics.
Phuong T. Nguyen 0001, Juri Di Rocco, Ludovico Iovino, Davide Di Ruscio, Alfonso Pierantonio
Softw. Syst. Model.2
2021 Correction to: Evaluation of a machine learning classifier for metamodels
Phuong T. Nguyen 0001, Juri Di Rocco, Ludovico Iovino, Davide Di Ruscio, Alfonso Pierantonio
Softw. Syst. Model.2
2020 TopFilter: An Approach to Recommend Relevant GitHub Topics
abstract
Background: In the context of software development, GitHub has been at the forefront of platforms to store, analyze and maintain a large number of software repositories. Topics have been introduced by GitHub as an effective method to annotate stored repositories. However, labeling GitHub repositories should be carefully conducted to avoid adverse effects on project popularity and reachability. Aims: We present TopFilter, a novel approach to assist open source software developers in selecting suitable topics for GitHub repositories being created. Method: We built a project-topic matrix and applied a syntactic-based similarity function to recommend missing topics by representing repositories and related topics in a graph. The ten-fold cross-validation methodology has been used to assess the performance of TopFilter by considering different metrics, i.e., success rate, precision, recall, and catalog coverage. Result: The results show that TopFilter recommends good topics depending on different factors, i.e., collaborative filtering settings, considered datasets, and pre-processing activities. Moreover, TopFilter can be combined with a state-of-the-art topic recommender system (i.e., MNB network) to improve the overall prediction performance. Conclusion: Our results confirm that collaborative filtering techniques can successfully be used to provide relevant topics for GitHub repositories. Moreover, TopFilter can gain a significant boost in prediction performances by employing the outcomes obtained by the MNB network as its initial set of topics.
Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen 0001, Riccardo Rubei
ESEM1
2020 Detecting Java software similarities by using different clustering techniques
Andrea Capiluppi, Davide Di Ruscio, Juri Di Rocco, Phuong T. Nguyen 0001, Nemitari Ajienka
Inf. Softw. Technol.3
2020 PostFinder: Mining Stack Overflow posts to support software developers
Riccardo Rubei, Claudio Di Sipio, Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio
Inf. Softw. Technol.4
2020 CrossRec: Supporting software developers by recommending third-party libraries
Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio, Massimiliano Di Penta
J. Syst. Softw.2
2020 Understanding MDE projects: megamodels to the rescue for architecture recovery
Juri Di Rocco, Davide Di Ruscio, Johannes Härtel, Ludovico Iovino, Ralf Lämmel, Alfonso Pierantonio
Softw. Syst. Model.1
2020 An automated approach to assess the similarity of GitHub repositories
Phuong T. Nguyen 0001, Juri Di Rocco, Riccardo Rubei, Davide Di Ruscio
Softw. Qual. J.2
2019 Enabling heterogeneous recommendations in OSS development: what's done and what's next in CROSSMINER
abstract
Open source software (OSS) forges contain rich data sources that are useful for supporting development activities. Research has been done to promote techniques and tools for providing open source developers with innovative features aiming at obtaining improvements in terms of development effort, cost savings, and developer productivity, just to mention a few. In the context of the EU H2020 CROSSMINER project we are conceiving a set of recommendations to assist software programmers in different phases of the development process. To this end, we defined a graph-based representation to encode in a homogeneous manner different aspects of OSS ecosystems as well as to incorporate various well-founded recommendation techniques. Following the proposed paradigm, we have implemented recommender systems for providing various artifacts, such as third-party libraries and API usage. The preliminary results we achieved so far are promising: our proposed systems are able to suggest highly relevant items with respect to the current development context. In this paper, we describe what has been achieved so far as well as our planned medium and longer-term objectives. As a proof of concept, we present a use case where we built a context-aware recommender system to recommend API function calls and usage patterns.
Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio
EASE2
2019 Query-Based Impact Analysis of Metamodel Evolutions
abstract
Metamodels are at the core of any modeling ecosystem. As their evolution is inevitable, the management of artifacts which depend on these metamodels is a complicated task. Restoring the validity of the corrupted artifacts after a metamodel evolution in a (semi-)automated manner is intrinsically difficult especially when considering the exact impact of the evolution on the restoring process. In this paper, we propose a generic approach to automatically quantify and identify the impact of metamodel evolution on two related artifacts: models and transformations. The approach starts from the evolution definition and generates OCL queries that can be executed on these artifacts to obtain the impacted elements. The knowledge gained from the impact analysis may then guide the user in the decision on whether to proceed with the evolution or to revert it.
Ludovico Iovino, Adrian Rutle, Alfonso Pierantonio, Juri Di Rocco
SEAA4
2019 FOCUS: a recommender system for mining API function calls and usage patterns
abstract
Software developers interact with APIs on a daily basis and, therefore, often face the need to learn how to use new APIs suitable for their purposes. Previous work has shown that recommending usage patterns to developers facilitates the learning process. Current approaches to usage pattern recommendation, however, still suffer from high redundancy and poor run-time performance. In this paper, we reformulate the problem of usage pattern recommendation in terms of a collaborative-filtering recommender system. We present a new tool, FOCUS, which mines open-source project repositories to recommend API method invocations and usage patterns by analyzing how APIs are used in projects similar to the current project. We evaluate FOCUS on a large number of Java projects extracted from GitHub and Maven Central and find that it outperforms the state-of-the-art approach PAM with regards to success rate, accuracy, and execution time. Results indicate the suitability of context-aware collaborative-filtering recommender systems to provide API usage patterns.
Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio, Lina Ochoa, Thomas Degueule, Massimiliano Di Penta
ICSE2
2019 Automated Classification of Metamodel Repositories: A Machine Learning Approach
abstract
Manual classification methods of metamodel repositories require highly trained personnel and the results are usually influenced by the subjectivity of human perception. Therefore, automated metamodel classification is very desirable and stringent. In this work, Machine Learning techniques have been employed for metamodel automated classification. In particular, a tool implementing a feed-forward neural network is introduced to classify metamodels. An experimental evaluation over a dataset of 555 metamodels demonstrates that the technique permits to learn from manually classified data and effectively categorize incoming unlabeled data with a considerably high prediction rate: the best performance comprehends 95.40% as success rate, 0.945 as precision, 0.938 as recall, and 0.942 as F1 score.
Phuong T. Nguyen 0001, Juri Di Rocco, Davide Di Ruscio, Alfonso Pierantonio, Ludovico Iovino
MoDELS2
2019 Automated Reuse of Model Transformations through Typing Requirements Models
abstract
Model transformations are key elements of model-driven engineering, where they are used to automate the manipulation of models. However, they are typed with respect to concrete source and target meta-models, making their reuse for other (even similar) meta-models challenging. To improve this situation, we propose capturing the typing requirements for reusing a transformation with other meta-models by the notion of a typing requirements model (TRM). A TRM describes the prerequisites that a model transformation imposes on the source and target meta-models to obtain a correct typing. The key observation is that any meta-model pair that satisfies the TRM is a valid reuse context for the transformation at hand. A TRM is made of two domain requirement models (DRMs) describing the requirements for the source and target meta-models, and a compatibility model expressing dependencies between them. We define a notion of refinement between DRMs and see meta-models as a special case of DRM. We provide a catalogue of valid refinements and describe how to automatically extract a TRM from an ATL transformation. The approach is supported by our tool TOTEM. We report on two experiments—based on transformations developed by third parties and meta-model mutation techniques—validating the correctness and completeness of our TRM extraction procedure and confirming the power of TRMs to encode variability and support flexible reuse.
Juan de Lara, Esther Guerra, Davide Di Ruscio, Juri Di Rocco, Jesús Sánchez Cuadrado, Ludovico Iovino, Alfonso Pierantonio
ACM Trans. Softw. Eng. Methodol.4
2018 CrossSim: Exploiting Mutual Relationships to Detect Similar OSS Projects
abstract
Software development is a knowledge-intensive activity, which requires mastering several languages, frameworks, technology trends (among other aspects) under the pressure of ever-increasing arrays of external libraries and resources. Recommender systems are gaining high relevance in software engineering since they aim at providing developers with real-time recommendations, which can reduce the time spent on discovering and understanding reusable artifacts from software repositories, and thus inducing productivity and quality gains. In this paper, we focus on the problem of mining open source software repositories to identify similar projects, which can be evaluated and eventually reused by developers. To this end, CrossSim is proposed as a novel approach to model open source software projects and related artifacts and to compute similarities among them. An evaluation on a dataset containing 580 GitHub projects shows that CrossSim outperforms an existing technique, which has been proven to have a good performance in detecting similar GitHub repositories.
Phuong T. Nguyen 0001, Juri Di Rocco, Riccardo Rubei, Davide Di Ruscio
SEAA2
2017 Reusing Model Transformations Through Typing Requirements Models
Juan de Lara, Juri Di Rocco, Davide Di Ruscio, Esther Guerra, Ludovico Iovino, Alfonso Pierantonio, Jesús Sánchez Cuadrado
FASE2
2016 Automated Clustering of Metamodel Repositories
Francesco Basciani, Juri Di Rocco, Davide Di Ruscio, Ludovico Iovino, Alfonso Pierantonio
CAiSE2
2015 Mining Correlations of ATL Model Transformation and Metamodel Metrics
abstract
Model transformations are considered to be the "heart" and "soul" of Model Driven Engineering, and as a such, advanced techniques and tools are needed for supporting the development, quality assurance, maintenance, and evolution of model transformations. Even though model transformation developers are gaining the availability of powerful languages and tools for developing, and testing model transformations, very few techniques are available to support the understanding of transformation characteristics. In this paper, we propose a process to analyze model transformations with the aim of identifying to what extent their characteristics depend on the corresponding input and target met models. The process relies on a number of transformation and metamodel metrics that are calculated and properly correlated. The paper discusses the application of the approach on a corpus consisting of more than 90 ATL transformations and 70 corresponding metamodels.
Juri Di Rocco, Davide Di Ruscio, Ludovico Iovino, Alfonso Pierantonio
MiSE@ICSE1
2015 Supporting users to manage breaking and unresolvable changes in coupled evolution
abstract
In Model-Driven Engineering (MDE) metamodels play a key role since they underpin the specification of different kinds of modeling artifacts, and the development of a wide range of model management tools. Consequently, when a metamodel is changed modelers and developers have to deal with the induced coupled evolutions i.e., adapting all those artifacts that might have been affected by the operated metamodel changes. Over the last years, several approaches have been proposed to deal with the coupled evolution problem, even though the treatment of changes is still a time consuming and error-prone activity. In this paper we propose an approach supporting users during the adaptation steps that cannot be fully automated.~The approach has been implemented by extending the EMFMigrate language and by exploiting the user input facility of the Epsilon Object Language. The approach has been applied to cope with the coupled evolution of metamodels and model-to-text transformations
Juri Di Rocco, Davide Di Ruscio, Alfonso Pierantonio, Ludovico Iovino
DSM@SPLASH1
2014 Mining metrics for understanding metamodel characteristics
abstract
Metamodels are a key concept in Model-Driven Engineering. Any artifact in a modeling ecosystem has to be defined in accordance to a metamodel prescribing its main qualities. Hence, understanding common characteristics of metamodels, how they evolve over time, and what is the impact of metamodel changes throughout the modeling ecosystem is of great relevance. Similarly to software, metrics can be used to obtain objective, transparent, and reproducible measurements on metamodels too. In this paper, we present an approach to understand structural characteristics of metamodels. A number of metrics are used to quantify and measure metamodels and cross-link different aspects in order to provide additional information about how metamodel characteristics are related. The approach is applied on repositories consisting of more than 450 metamodels.
Juri Di Rocco, Davide Di Ruscio, Ludovico Iovino, Alfonso Pierantonio
MiSE1
2014 Models of OSS project meta-information: a dataset of three forges
abstract
The process of selecting open-source software (OSS) for adoption is not straightforward as it involves exploring various sources of information to determine the quality, maturity, activity, and user support of each project. In the context of the OSSMETER project, we have developed a forge-agnostic metamodel that captures the meta-information common to all OSS projects. We specialise this metamodel for popular OSS forges in order to capture forge-specific meta-information. In this paper we present a dataset conforming to these metamodels for over 500,000 OSS projects hosted on three popular OSS forges: Eclipse, SourceForge, and GitHub. The dataset enables different kinds of automatic analysis and supports objective comparisons of cross-forge OSS alternatives with respect to a user's needs and quality requirements.
James R. Williams, Davide Di Ruscio, Nicholas Drivalos Matragkas, Juri Di Rocco, Dimitrios S. Kolovos
MSR4