VLDB 2026 Research / reviewers in the wild / expert
Luigi Quaranta
dblp:238/0016
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-9221-0739ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-monitoring of Developers' Emotions: The Case of Agile Retrospective MeetingsabstractDevelopers experience a wide range of emotions while creating software. Being able to identify the causes of one’s own and peers’ emotions can equip developers with the ability to regulate their behavior to restore positive moods and productivity. In this article, we investigate to what extent self-monitoring of emotions can enhance agile retrospective meetings by improving the emotion awareness of participants. To this aim, we conducted a controlled experiment involving three software development teams involving two student teams and one professional developers team. The experimental design involves the collection of biometrics and self-reported information about emotions, which are then visualized before the retrospective meetings to inform discussion using EmoVizPhy, a tool that we designed and implemented for this aim. While students found that self-monitoring helped them recall significant emotional episodes, leading to more meaningful contributions during retrospectives, professional developers perceived limited benefits from this practice. Furthermore, based on the analysis of corrective actions identified by the participants during the study, we hypothesize that self-monitoring of emotions through EmoVizPhy may play a valuable role in facilitating the consolidation of new agile teams for which roles and collaboration dynamics are still being defined. Daniela Grassi, Filippo Lanubile, Nicole Novielli, Luigi Quaranta, Alexander Serebrenik |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | MLOps in the Healthcare Domain: a Systematic Literature Review
Giulio Mallardi, Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
SEAA (2) | 2 |
| 2025 | A multivocal literature review on the benefits and limitations of industry-leading AutoML tools
Luigi Quaranta, Kelly Azevedo, Fabio Calefato, Marcos Kalinowski |
Inf. Softw. Technol. | 1 |
| 2024 | An MLOps Approach for Deploying Machine Learning Models in Healthcare SystemsabstractIn recent years, there has been a remarkable increase in the use of machine learning (ML) technologies in healthcare settings. Despite this growth, a significant challenge persists: numerous promising initiatives remain confined to research laboratories, unable to make the critical transition into clinical practice. While the gap between research and production deployment affects ML projects across various sectors, the stringently regulated healthcare environment poses unique and heightened challenges. To address these challenges, MLOps has recently emerged as a specialized discipline that combines engineering best practices with operational excellence. Building upon software engineering foundations and DevOps principles, MLOps introduces a systematic approach to automating ML workflows and managing the complete model lifecycle. This paper introduces a practical and comprehensive MLOps-based framework. This framework is designed to facilitate the transformation of experimental ML models into production-ready healthcare solutions. It provides a structured approach that ensures the seamless integration of ML-powered tools into clinical environments and guarantees their reliability and compliance with medical standards, instilling confidence in their effectiveness. We are currently implementing and evaluating this framework within the "DARE – Digital Lifelong Prevention" project, a national Italian initiative aiming to harness data analytics to enhance preventive healthcare strategies across different life stages. Giulio Mallardi, Fabio Calefato, Luigi Quaranta, Filippo Lanubile |
BIBM | 3 |
| 2024 | Continuous Quality Improvement of AI-based Systems: the QualAI ProjectabstractQualAI is a two-year project aimed at defining a set of recommenders to continuously monitor, assess, and improve the quality of AI-based systems, with a particular focus on machine learning (ML) applications. We will develop recommenders for the quality assurance of both data and ML models to enable practitioners to mitigate technical debt. Special attention will be paid to communication challenges that may arise in hybrid teams comprising data scientists and software developers. This paper presents the project outline, provides an executive summary of the research activities, outlines the expected project outcomes, and reports the results obtained to date. Nicole Novielli, Rocco Oliveto, Fabio Palomba, Fabio Calefato, Giuseppe Colavito, Vincenzo De Martino, Antonio Della Porta, Giammaria Giordano, Emanuela Guglielmi, Filippo Lanubile, Luigi Quaranta, Gilberto Recupito, Simone Scalabrino, Angelica Spina, Antonio Vitale |
ESEM | 11 |
| 2024 | Leveraging GPT-like LLMs to Automate Issue LabelingabstractIssue labeling is a crucial task for the effective management of software projects. To date, several approaches have been put forth for the automatic assignment of labels to issue reports. In particular, supervised approaches based on the fine-tuning of BERT-like language models have been proposed, achieving state-of-the-art performance. More recently, decoder-only models such as GPT have become prominent in SE research due to their surprising capabilities to achieve state-of-the-art performance even for tasks they have not been trained for. To the best of our knowledge, GPT-like models have not been applied yet to the problem of issue classification, despite the promising results achieved for many other software engineering tasks. In this paper, we investigate to what extent we can leverage GPT-like LLMs to automate the issue labeling task. Our results demonstrate the ability of GPT-like models to correctly classify issue reports in the absence of labeled data that would be required to fine-tune BERT-like LLMs. Giuseppe Colavito, Filippo Lanubile, Nicole Novielli, Luigi Quaranta |
MSR | 4 |
| 2024 | A lot of talk and a badge: An exploratory analysis of personal achievements in GitHubabstractContext: GitHub has introduced a new gamification element through personal achievements, whereby badges are unlocked and displayed on developers’ personal profile pages in recognition of their development activities. Objective: In this paper, we present an exploratory analysis using mixed methods to study the diffusion of personal badges in GitHub, in addition to the effects and reactions to their introduction. Method: First, we conduct an observational study by mining longitudinal data from more than 6,000 developers and performed correlation and regression analysis. Then, we conduct a survey and analyze over 300 GitHub community discussions on the topic of personal badges to gauge how the community responded to the introduction of the new feature. Results: We find that most of the developers sampled own at least a badge, but we also observe an increasing number of users who choose to keep their profile private and opt out of displaying badges. Additionally, badges are generally poorly correlated with developers’ skills and dispositions such as timeliness and desire to collaborate. We also find that, except for the Starstruck badge (reflecting the number of followers), their introduction does not have an effect. Finally, the reaction of the community has been in general mixed, as developers find them appealing in principle but without a clear purpose and hardly reflecting their abilities in the current form. Conclusions: We provide recommendations to the designers of the GitHubplatform on how to improve the current implementation of personal badges as both a gamification mechanism and as sources of reliable cues for assessing the abilities of developers. Fabio Calefato, Luigi Quaranta, Filippo Lanubile |
Inf. Softw. Technol. | 2 |
| 2024 | Impact of data quality for automatic issue classification using pre-trained language models
Giuseppe Colavito, Filippo Lanubile, Nicole Novielli, Luigi Quaranta |
J. Syst. Softw. | 4 |
| 2023 | Assessing the Use of AutoML for Data-Driven Software EngineeringabstractBackground. Due to the widespread adoption of Artificial Intelligence (AI) and Machine Learning (ML) for building software applications, companies are struggling to recruit employees with a deep understanding of such technologies. In this scenario, AutoML is soaring as a promising solution to fill the AI/ML skills gap since it promises to automate the building of end-to-end AI/ML pipelines that would normally be engineered by specialized team members. Aims. Despite the growing interest and high expectations, there is a dearth of information about the extent to which AutoML is currently adopted by teams developing AI/ML-enabled systems and how it is perceived by practitioners and researchers. Method. To fill these gaps, in this paper, we present a mixed-method study comprising a benchmark of 12 end-to-end AutoML tools on two SE datasets and a user survey with follow-up interviews to further our understanding of AutoML adoption and perception. Results. We found that AutoML solutions can generate models that outperform those trained and optimized by researchers to perform classification tasks in the SE domain. Also, our findings show that the currently available AutoML solutions do not live up to their names as they do not equally support automation across the stages of the ML development workflow and for all the team members. Conclusions. We derive insights to inform the SE research community on how AutoML can facilitate their activities and tool builders on how to design the next generation of AutoML technologies. Fabio Calefato, Luigi Quaranta, Filippo Lanubile, Marcos Kalinowski |
ESEM | 2 |
| 2022 | Pynblint: a static analyzer for Python Jupyter notebooksabstractJupyter Notebook is the tool of choice of many data scientists in the early stages of ML workflows. The notebook format, however, has been criticized for inducing bad programming practices; indeed, researchers have already shown that open-source repositories are inundated by poor-quality notebooks. Low-quality output from the prototypical stages of ML workflows constitutes a clear bottleneck towards the productization of ML models. To foster the creation of better notebooks, we developed Pynblint, a static analyzer for Jupyter notebooks written in Python. The tool checks the compliance of notebooks (and surrounding repositories) with a set of empirically validated best practices and provides targeted recommendations when violations are detected. Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
CAIN | 1 |
| 2022 | A Preliminary Investigation of MLOps Practices in GitHubabstractBackground. The rapid and growing popularity of machine learning (ML) applications has led to an increasing interest in MLOps, that is, the practice of continuous integration and deployment (CI/CD) of ML-enabled systems. Aims. Since changes may affect not only the code but also the ML model parameters and the data themselves, the automation of traditional CI/CD needs to be extended to manage model retraining in production. Method. In this paper, we present an initial investigation of the MLOps practices implemented in a set of ML-enabled systems retrieved from GitHub, focusing on GitHub Actions and CML, two solutions to automate the development workflow. Results. Our preliminary results suggest that the adoption of MLOps workflows in open-source GitHub projects is currently rather limited. Conclusions. Issues are also identified, which can guide future research work. Fabio Calefato, Filippo Lanubile, Luigi Quaranta |
ESEM | 3 |
| 2022 | Eliciting Best Practices for Collaboration with Computational NotebooksabstractDespite the widespread adoption of computational notebooks, little is known about best practices for their usage in collaborative contexts. In this paper, we fill this gap by eliciting a catalog of best practices for collaborative data science with computational notebooks. With this aim, we first look for best practices through a multivocal literature review. Then, we conduct interviews with professional data scientists to assess their awareness of these best practices. Finally, we assess the adoption of best practices through the analysis of 1,380 Jupyter notebooks retrieved from the Kaggle platform. Findings reveal that experts are mostly aware of the best practices and tend to adopt them in their daily work. Nonetheless, they do not consistently follow all the recommendations as, depending on specific contexts, some are deemed unfeasible or counterproductive due to the lack of proper tool support. As such, we envision the design of notebook solutions that allow data scientists not to have to prioritize exploration and rapid prototyping over writing code of quality. Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | KGTorrent: A Dataset of Python Jupyter Notebooks from KaggleabstractComputational notebooks have become the tool of choice for many data scientists and practitioners for performing analyses and disseminating results. Despite their increasing popularity, the research community cannot yet count on a large, curated dataset of computational notebooks. In this paper, we fill this gap by introducing KGTorrent, a dataset of Python Jupyter notebooks with rich metadata retrieved from Kaggle, a platform hosting data science competitions for learners and practitioners with any levels of expertise. We describe how we built KGTorrent, and provide instructions on how to use it and refresh the collection to keep it up to date. Our vision is that the research community will use KGTorrent to study how data scientists, especially practitioners, use Jupyter Notebook in the wild and identify potential shortcomings to inform the design of its future extensions. Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
MSR | 1 |
| 2019 | A replication study on code comprehension and expertise using lightweight biometric sensorsabstractCode comprehension has been recently investigated from physiological and cognitive perspectives using medical imaging devices. Floyd et al. (i.e., the original study) used fMRI to classify the type of comprehension tasks performed by developers and relate their results to their expertise. We replicate the original study using lightweight biometrics sensors. Our study participants-28 undergrads in computer science-performed comprehension tasks on source code and natural language prose. We developed machine learning models to automatically identify what kind of tasks developers are working on leveraging their brain-, heart-, and skin-related signals. The best improvement over the original study performance is achieved using solely the heart signal obtained through a single device (BAC 87%vs. 79.1%). Differently from the original study, we did not observe a correlation between the participants' expertise and the classifier performance (τ= 0.16, p= 0.31). Our findings show that lightweight biometric sensors can be used to accurately recognize comprehension opening interesting scenarios for research and practice. Davide Fucci, Daniela Girardi, Nicole Novielli, Luigi Quaranta, Filippo Lanubile |
ICPC | 4 |