VLDB 2026 Research / reviewers in the wild / expert
Fabio Calefato
dblp:48/5662
· DBLP profile ↗
56ranked-venue papers
38as first author
19since 2021 · last 2026
0000-0003-2654-1588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 46 · 31 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Who "controls" where work shall be done? State-of-practice in post-pandemic remote work regulationabstractABSTRACT The COVID-19 pandemic has permanently altered workplace structures, making remote work a widespread practice. While many employees advocate for flexibility, many employers reconsider their attitude toward remote work and opt for structured return-to-office mandates. Media headlines repeatedly emphasize that the corporate world returns to full-time office work. This study examines how companies in software-intensive industry regulate work location, whether corporate policies have evolved in the last five years, and, if so, how, and why. We collected data on remote work regulation from corporate HR and management representatives from 68 companies that vary in size, location, and preferred work modality. Our findings reveal that although many companies prioritize office-oriented work (50%), most companies in our sample permit hybrid work (84%) and only four companies are returning to full-time office work. Remote work regulation does not reveal any particular new “best practice” as policies differ greatly; however, the single most popular arrangement was the three in-office days per week. More than half of the companies (53%) encourage or mandate office attendance centrally, with additional 18% having decentralized mandates. Over a quarter (28%) have changed regulations gradually increasing the mandatory office presence or implementing differentiated conditions. Our key recommendation for office-oriented companies is to consider trust-based recommendations as an alternative to centralized office presence mandates, while for companies oriented toward remote working, we warn about the points of no (or hard) return. Finally, the current state of policies is clearly not final, as companies continue to experiment and adjust their work regulation Darja Smite, Nils Brede Moe, Maria Teresa Baldassarre, Fabio Calefato, Guilherme Horta Travassos, Marcin Floryan, Marcos Kalinowski, Daniel Méndez 0001, Graziela Pereira, Margaret-Anne D. Storey, Rafael Prikladnicki |
J. Syst. Softw. | 4 |
| 2026 | Self-Admitted GenAI Usage in Open-Source SoftwareabstractThe widespread adoption of generative AI (GenAI) tools such as GitHub Copilot and ChatGPT is transforming software development. Since generated source code is virtually impossible to distinguish from manually written code, their real-world usage and impact on opensource software (OSS) development remain poorly understood. In this paper, we introduce the concept of self-admitted GenAI usage, that is, developers explicitly referring to the use of GenAI tools for content creation in software artifacts. Using this concept as a lens to study how GenAI tools are integrated into OSS projects, we analyze a curated sample of more than 200,000 GitHub repositories, identifying 1,292 such self-admissions across 156 repositories in commit messages, code comments, and project documentation. Using a mixed methods approach, we derive a taxonomy of 32 tasks, 10 content types, and 11 purposes associated with GenAI usage based on 1,292 qualitatively coded mentions. We then analyze 13 documents with policies and usage guidelines for GenAI tools and conduct a developer survey to uncover the ethical, legal, and practical concerns behind them. Our findings reveal that developers actively manage how GenAI is used in their projects, highlighting the need for project-level transparency, attribution, and quality control practices in AI-assisted software development. Finally, we examine the longitudinal impact of GenAI adoption on code churn in 151 repositories with self-admitted GenAI usage and find no general increase, contradicting popular narratives on the impact of GenAI on software development. Tao Xiao 0001, Youmei Fan, Fabio Calefato, Christoph Treude, Raula Gaikovina Kula, Hideaki Hata, Sebastian Baltes |
IEEE Trans. Software Eng. | 3 |
| 2025 | Exploring Engagement in Hybrid MeetingsabstractBackground. The widespread adoption of hybrid work following the COVID-19 pandemic has fundamentally transformed software development practices, introducing new challenges in communication and collaboration as organizations transition from traditional office-based structures to flexible working arrangements. This shift has established a new organizational norm where even traditionally office-first companies now embrace hybrid team structures. While remote participation in meetings has become commonplace in this new environment, it may lead to isolation, alienation, and decreased engagement among remote team members. Aims. This study aims to identify and characterize engagement patterns in hybrid meetings through objective measurements, focusing on the differences between co-located and remote participants. Method. We studied professionals from three software companies over several weeks, employing a multimodal approach to measure engagement. Data were collected through self-reported questionnaires and physiological measurements using biometric devices during hybrid meetings to understand engagement dynamics. Results. The regression analyses revealed comparable engagement levels between onsite and remote participants, though remote participants show lower engagement in long meetings regardless of participation mode. Active roles positively correlate with higher engagement, while larger meetings and afternoon sessions are associated with lower engagement. Conclusions. Our results offer insights into factors associated with engagement and disengagement in hybrid meetings, as well as potential meeting improvement recommendations. These insights are potentially relevant not only for software teams but also for knowledge-intensive organizations across various sectors facing similar hybrid collaboration challenges. Daniela Grassi, Fabio Calefato, Darja Smite, Nicole Novielli, Filippo Lanubile |
ESEM | 2 |
| 2025 | MLOps in the Healthcare Domain: a Systematic Literature Review
Giulio Mallardi, Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
SEAA (2) | 3 |
| 2025 | A multivocal literature review on the benefits and limitations of industry-leading AutoML tools
Luigi Quaranta, Kelly Azevedo, Fabio Calefato, Marcos Kalinowski |
Inf. Softw. Technol. | 3 |
| 2024 | An MLOps Approach for Deploying Machine Learning Models in Healthcare SystemsabstractIn recent years, there has been a remarkable increase in the use of machine learning (ML) technologies in healthcare settings. Despite this growth, a significant challenge persists: numerous promising initiatives remain confined to research laboratories, unable to make the critical transition into clinical practice. While the gap between research and production deployment affects ML projects across various sectors, the stringently regulated healthcare environment poses unique and heightened challenges. To address these challenges, MLOps has recently emerged as a specialized discipline that combines engineering best practices with operational excellence. Building upon software engineering foundations and DevOps principles, MLOps introduces a systematic approach to automating ML workflows and managing the complete model lifecycle. This paper introduces a practical and comprehensive MLOps-based framework. This framework is designed to facilitate the transformation of experimental ML models into production-ready healthcare solutions. It provides a structured approach that ensures the seamless integration of ML-powered tools into clinical environments and guarantees their reliability and compliance with medical standards, instilling confidence in their effectiveness. We are currently implementing and evaluating this framework within the "DARE – Digital Lifelong Prevention" project, a national Italian initiative aiming to harness data analytics to enhance preventive healthcare strategies across different life stages. Giulio Mallardi, Fabio Calefato, Luigi Quaranta, Filippo Lanubile |
BIBM | 2 |
| 2024 | Continuous Quality Improvement of AI-based Systems: the QualAI ProjectabstractQualAI is a two-year project aimed at defining a set of recommenders to continuously monitor, assess, and improve the quality of AI-based systems, with a particular focus on machine learning (ML) applications. We will develop recommenders for the quality assurance of both data and ML models to enable practitioners to mitigate technical debt. Special attention will be paid to communication challenges that may arise in hybrid teams comprising data scientists and software developers. This paper presents the project outline, provides an executive summary of the research activities, outlines the expected project outcomes, and reports the results obtained to date. Nicole Novielli, Rocco Oliveto, Fabio Palomba, Fabio Calefato, Giuseppe Colavito, Vincenzo De Martino, Antonio Della Porta, Giammaria Giordano, Emanuela Guglielmi, Filippo Lanubile, Luigi Quaranta, Gilberto Recupito, Simone Scalabrino, Angelica Spina, Antonio Vitale |
ESEM | 4 |
| 2024 | A lot of talk and a badge: An exploratory analysis of personal achievements in GitHubabstractContext: GitHub has introduced a new gamification element through personal achievements, whereby badges are unlocked and displayed on developers’ personal profile pages in recognition of their development activities. Objective: In this paper, we present an exploratory analysis using mixed methods to study the diffusion of personal badges in GitHub, in addition to the effects and reactions to their introduction. Method: First, we conduct an observational study by mining longitudinal data from more than 6,000 developers and performed correlation and regression analysis. Then, we conduct a survey and analyze over 300 GitHub community discussions on the topic of personal badges to gauge how the community responded to the introduction of the new feature. Results: We find that most of the developers sampled own at least a badge, but we also observe an increasing number of users who choose to keep their profile private and opt out of displaying badges. Additionally, badges are generally poorly correlated with developers’ skills and dispositions such as timeliness and desire to collaborate. We also find that, except for the Starstruck badge (reflecting the number of followers), their introduction does not have an effect. Finally, the reaction of the community has been in general mixed, as developers find them appealing in principle but without a clear purpose and hardly reflecting their abilities in the current form. Conclusions: We provide recommendations to the designers of the GitHubplatform on how to improve the current implementation of personal badges as both a gamification mechanism and as sources of reliable cues for assessing the abilities of developers. Fabio Calefato, Luigi Quaranta, Filippo Lanubile |
Inf. Softw. Technol. | 1 |
| 2024 | Generative AI in Software Engineering Must Be Human-Centered: The Copenhagen Manifesto
Daniel Russo 0002, Sebastian Baltes, Niels van Berkel, Paris Avgeriou, Fabio Calefato, Beatriz Cabrero-Daniel, Gemma Catolino, Jürgen Cito, Neil A. Ernst, Thomas Fritz 0001, Hideaki Hata, Reid Holmes, Maliheh Izadi, Foutse Khomh, Mikkel Baun Kjærgaard, Grischa Liebel, Alberto Lluch-Lafuente, Stefano Lambiase, Walid Maalej, Gail C. Murphy, Nils Brede Moe, Gabrielle O'Brien, Elda Paja, Mauro Pezzè, John Stouby Persson, Rafael Prikladnicki, Paul Ralph, Martin P. Robillard, Thiago Rocha Silva, Klaas-Jan Stol, Margaret-Anne D. Storey, Viktoria Stray, Paolo Tell, Christoph Treude, Bogdan Vasilescu |
J. Syst. Softw. | 5 |
| 2023 | Assessing the Use of AutoML for Data-Driven Software EngineeringabstractBackground. Due to the widespread adoption of Artificial Intelligence (AI) and Machine Learning (ML) for building software applications, companies are struggling to recruit employees with a deep understanding of such technologies. In this scenario, AutoML is soaring as a promising solution to fill the AI/ML skills gap since it promises to automate the building of end-to-end AI/ML pipelines that would normally be engineered by specialized team members. Aims. Despite the growing interest and high expectations, there is a dearth of information about the extent to which AutoML is currently adopted by teams developing AI/ML-enabled systems and how it is perceived by practitioners and researchers. Method. To fill these gaps, in this paper, we present a mixed-method study comprising a benchmark of 12 end-to-end AutoML tools on two SE datasets and a user survey with follow-up interviews to further our understanding of AutoML adoption and perception. Results. We found that AutoML solutions can generate models that outperform those trained and optimized by researchers to perform classification tasks in the SE domain. Also, our findings show that the currently available AutoML solutions do not live up to their names as they do not equally support automation across the stages of the ML development workflow and for all the team members. Conclusions. We derive insights to inform the SE research community on how AutoML can facilitate their activities and tool builders on how to design the next generation of AutoML technologies. Fabio Calefato, Luigi Quaranta, Filippo Lanubile, Marcos Kalinowski |
ESEM | 1 |
| 2022 | Pynblint: a static analyzer for Python Jupyter notebooksabstractJupyter Notebook is the tool of choice of many data scientists in the early stages of ML workflows. The notebook format, however, has been criticized for inducing bad programming practices; indeed, researchers have already shown that open-source repositories are inundated by poor-quality notebooks. Low-quality output from the prototypical stages of ML workflows constitutes a clear bottleneck towards the productization of ML models. To foster the creation of better notebooks, we developed Pynblint, a static analyzer for Jupyter notebooks written in Python. The tool checks the compliance of notebooks (and surrounding repositories) with a set of empirically validated best practices and provides targeted recommendations when violations are detected. Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
CAIN | 2 |
| 2022 | A Preliminary Investigation of MLOps Practices in GitHubabstractBackground. The rapid and growing popularity of machine learning (ML) applications has led to an increasing interest in MLOps, that is, the practice of continuous integration and deployment (CI/CD) of ML-enabled systems. Aims. Since changes may affect not only the code but also the ML model parameters and the data themselves, the automation of traditional CI/CD needs to be extended to manage model retraining in production. Method. In this paper, we present an initial investigation of the MLOps practices implemented in a set of ML-enabled systems retrieved from GitHub, focusing on GitHub Actions and CML, two solutions to automate the development workflow. Results. Our preliminary results suggest that the adoption of MLOps workflows in open-source GitHub projects is currently rather limited. Conclusions. Issues are also identified, which can guide future research work. Fabio Calefato, Filippo Lanubile, Luigi Quaranta |
ESEM | 1 |
| 2022 | Will you come back to contribute? Investigating the inactivity of OSS core developers in GitHubabstractAbstract Several Open-Source Software (OSS) projects depend on the continuity of their development communities to remain sustainable. Understanding how developers become inactive or why they take breaks can help communities prevent abandonment and incentivize developers to come back. In this paper, we propose a novel method to identify developers’ inactive periods by analyzing the individual rhythm of contributions to the projects. Using this method, we quantitatively analyze the inactivity of core developers in 18 OSS organizations hosted on GitHub. We also survey core developers to receive their feedback about the identified breaks and transitions. Our results show that our method was effective for identifying developers’ breaks. About 94% of the surveyed core developers agreed with our state model of inactivity; 71% and 79% of them acknowledged their breaks and state transition, respectively. We also show that all core developers take breaks (at least once) and about a half of them (~45%) have completely disengaged from a project for at least one year. We also analyzed the probability of transitions to/from inactivity and found that developers who pause their activity have a ~35 to ~55% chance to return to an active state; yet, if the break lasts for a year or longer, then the probability of resuming activities drops to ~21–26%, with a ~54% chance of complete disengagement. These results may support the creation of policies and mechanisms to make OSS community managers aware of breaks and potential project abandonment. Fabio Calefato, Marco Aurélio Gerosa, Giuseppe Iaffaldano, Filippo Lanubile, Igor Steinmacher |
Empir. Softw. Eng. | 1 |
| 2022 | Eliciting Best Practices for Collaboration with Computational NotebooksabstractDespite the widespread adoption of computational notebooks, little is known about best practices for their usage in collaborative contexts. In this paper, we fill this gap by eliciting a catalog of best practices for collaborative data science with computational notebooks. With this aim, we first look for best practices through a multivocal literature review. Then, we conduct interviews with professional data scientists to assess their awareness of these best practices. Finally, we assess the adoption of best practices through the analysis of 1,380 Jupyter notebooks retrieved from the Kaggle platform. Findings reveal that experts are mostly aware of the best practices and tend to adopt them in their daily work. Nonetheless, they do not consistently follow all the recommendations as, depending on specific contexts, some are deemed unfeasible or counterproductive due to the lack of proper tool support. As such, we envision the design of notebook solutions that allow data scientists not to have to prioritize exploration and rapid prototyping over writing code of quality. Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Using Personality Detection Tools for Software Engineering Research: How Far Can We Go?abstractAssessing the personality of software engineers may help to match individual traits with the characteristics of development activities such as code review and testing, as well as support managers in team composition. However, self-assessment questionnaires are not a practical solution for collecting multiple observations on a large scale. Instead, automatic personality detection, while overcoming these limitations, is based on off-the-shelf solutions trained on non-technical corpora, which might not be readily applicable to technical domains like software engineering. In this article, we first assess the performance of general-purpose personality detection tools when applied to a technical corpus of developers’ e-mails retrieved from the public archives of the Apache Software Foundation. We observe a general low accuracy of predictions and an overall disagreement among the tools. Second, we replicate two previous research studies in software engineering by replacing the personality detection tool used to infer developers’ personalities from pull-request discussions and e-mails. We observe that the original results are not confirmed, i.e., changing the tool used in the original study leads to diverging conclusions. Our results suggest a need for personality detection tools specially targeted for the software engineering domain. Fabio Calefato, Filippo Lanubile |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2022 | What Makes Agile Software Development Agile?abstractTogether with many success stories, promises such as the increase in production speed and the improvement in stakeholders’ collaboration have contributed to making agile a transformation in the software industry in which many companies want to take part. However, driven either by a natural and expected evolution or by contextual factors that challenge the adoption of agile methods as prescribed by their creator(s), software processes in practice mutate into hybrids over time. Are these still agile? In this article, we investigate the question: what makes a software development method agile? We present an empirical study grounded in a large-scale international survey that aims to identify software development methods and practices that improve or tame agility. Based on 556 data points, we analyze the perceived degree of agility in the implementation of standard project disciplines and its relation to used development methods and practices. Our findings suggest that only a small number of participants operate their projects in a purely traditional or agile manner (under 15 percent). That said, most project disciplines and most practices show a clear trend towards increasing degrees of agility. Compared to the methods used to develop software, the selection of practices has a stronger effect on the degree of agility of a given discipline. Finally, there are no methods or practices that explicitly guarantee or prevent agility. We conclude that agility cannot be defined solely at the process level. Additional factors need to be taken into account when trying to implement or improve agility in a software company. Finally, we discuss the field of software process-related research in the light of our findings and present a roadmap for future research. Marco Kuhrmann, Paolo Tell, Regina Hebig, Jil Klünder, Jürgen Münch, Oliver Linssen, Dietmar Pfahl, Michael Felderer, Christian Prause, Stephen G. MacDonell, Joyce Nakatumba-Nabende, David Raffo, Sarah Beecham, Eray Tüzün, Gustavo López 0001, Nicolás Paez, Diego Fontdevila, Sherlock A. Licorish, Steffen Küpper, Günther Ruhe, Eric Knauss, Özden Özcan Top, Paul M. Clarke, Fergal McCaffery, Marcela Genero, Aurora Vizcaíno, Mario Piattini, Marcos Kalinowski, Tayana Conte, Rafael Prikladnicki, Stephan Krusche, Ahmet Coskunçay, Ezequiel Scott, Fabio Calefato, Svetlana Pimonova, Rolf-Helge Pfeiffer, Ulrik Pagh Schultz Lundquist, Rogardt Heldal, Masud Fazal-Baqaie, Craig Anslow, Maleknaz Nayebi, Kurt Schneider, Stefan Sauer 0001, Dietmar Winkler 0001, Stefan Biffl, M. Cecilia Bastarrica, Ita Richardson |
IEEE Trans. Software Eng. | 34 |
| 2021 | KGTorrent: A Dataset of Python Jupyter Notebooks from KaggleabstractComputational notebooks have become the tool of choice for many data scientists and practitioners for performing analyses and disseminating results. Despite their increasing popularity, the research community cannot yet count on a large, curated dataset of computational notebooks. In this paper, we fill this gap by introducing KGTorrent, a dataset of Python Jupyter notebooks with rich metadata retrieved from Kaggle, a platform hosting data science competitions for learners and practitioners with any levels of expertise. We describe how we built KGTorrent, and provide instructions on how to use it and refresh the collection to keep it up to date. Our vision is that the research community will use KGTorrent to study how data scientists, especially practitioners, use Jupyter Notebook in the wild and identify potential shortcomings to inform the design of its future extensions. Luigi Quaranta, Fabio Calefato, Filippo Lanubile |
MSR | 2 |
| 2021 | Assessment of off-the-shelf SE-specific sentiment analysis tools: An extended replication studyabstractAbstract Sentiment analysis methods have become popular for investigating human communication, including discussions related to software projects. Since general-purpose sentiment analysis tools do not fit well with the information exchanged by software developers, new tools, specific for software engineering (SE), have been developed. We investigate to what extent off-the-shelf SE-specific tools for sentiment analysis mitigate the threats to conclusion validity of empirical studies in software engineering, highlighted by previous research. First, we replicate two studies addressing the role of sentiment in security discussions on GitHub and in question-writing on Stack Overflow. Then, we extend the previous studies by assessing to what extent the tools agree with each other and with the manual annotation on a gold standard of 600 documents. We find that different SE-specific sentiment analysis tools might lead to contradictory results at a fine-grain level, when used off-the-shelf. Conversely, platform-specific tuning or retraining might be needed to take into account differences in platform conventions, jargon, or document lengths. Nicole Novielli, Fabio Calefato, Filippo Lanubile, Alexander Serebrenik |
Empir. Softw. Eng. | 2 |
| 2021 | Global Software Engineering: Challenges and solutions
Fabio Calefato, Alpana Dubey, Christof Ebert, Paolo Tell |
J. Syst. Softw. | 1 |
| 2020 | A case study on tool support for collaboration in agile developmentabstractWe report on a longitudinal case study conducted at the Italian site of a large software company to further our understanding of how development and communication tools can be improved to better support agile practices and collaboration. After observing inconsistencies in the way communication tools (i.e., email, Skype, and Slack) were used, we first reinforced the use of Slack as the central hub for internal communication, while setting clear rules regarding tools usage. As a second main change, we refactored the Jira Scrum board into two separate boards, a detailed one for developers and a high-level one for managers, while also introducing automation rules and the integration with Slack. The first change revealed that the teams of developers used and appreciated Slack differently with the QA team being the most favorable and that the use of channels is hindered by automatic notifications from development tools (e.g., Jenkins). The findings from the second change show that 85% of the interviewees reported perceived improvements in their workflow. Despite the limitations due to the single nature of the reported case, we highlight the importance for companies to reflect on how to properly set up their agile work environment to improve communication and facilitate collaboration. Fabio Calefato, Andrea Giove, Filippo Lanubile, Marco Losavio |
ICGSE | 1 |
| 2020 | Can We Use SE-specific Sentiment Analysis Tools in a Cross-Platform Setting?abstractIn this paper, we address the problem of using sentiment analysis tools 'off-the-shelf', that is when a gold standard is not available for retraining. We evaluate the performance of four SE-specific tools in a cross-platform setting, i.e., on a test set collected from data sources different from the one used for training. We find that (i) the lexicon-based tools outperform the supervised approaches retrained in a cross-platform setting and (ii) retraining can be beneficial in within-platform settings in the presence of robust gold standard datasets, even using a minimal training set. Based on our empirical findings, we derive guidelines for reliable use of sentiment analysis tools in software engineering. Nicole Novielli, Fabio Calefato, Davide Dongiovanni, Daniela Girardi, Filippo Lanubile |
MSR | 2 |
| 2020 | The Impact of Dynamics of Collaborative Software Engineering on Introverts: A Study ProtocolabstractBackground: Collaboration among software engineers through face-to-face discussions in teams has been promoted since the adoption of agile methods. However, these discussions might demote the contribution of software engineers who are introverts, possibly leading to sub-optimal solutions and creating work environments that benefit extroverts. Objective: We aim to evaluate whether providing software engineers with time to work individually and reason about a collective problem is a setting that makes introverts more comfortable to interact and contribute more, ultimately leading to better solutions. Method: We plan to conduct a between-subjects study, with teams in a control group that design a software architecture in a team discussion meeting and teams in a treatment group in which subjects work individually before engaging in a meeting. We will assess and compare the amount of contribution of introverts, their subjective experiences, and the designed solutions. Limitations: As extroverts will be present in both groups, we will not be able to conclude that better solutions are solely due to the increased participation of introverts. The analyses of their subjective experience and amount of contributions might provide evidence to suggest the reasons for observed differences. Ingrid Nunes, Christoph Treude, Fabio Calefato |
MSR | 3 |
| 2019 | An empirical assessment of best-answer prediction models in technical Q&A sites
Fabio Calefato, Filippo Lanubile, Nicole Novielli |
Empir. Softw. Eng. | 1 |
| 2019 | A large-scale, in-depth analysis of developers' personalities in the Apache ecosystem
Fabio Calefato, Filippo Lanubile, Bogdan Vasilescu |
Inf. Softw. Technol. | 1 |
| 2019 | RECODE: revision control for digital images
Fabio Calefato, Giovanna Castellano, Veronica Rossano |
Multim. Tools Appl. | 1 |
| 2019 | Correction to: RECODE: revision control for digital images
Fabio Calefato, Giovanna Castellano, Veronica Rossano |
Multim. Tools Appl. | 1 |
| 2018 | Collaboration Success Factors in an Online Music CommunityabstractOnline communities have been able to develop large, open-source software (OSS) projects like Linux and Firefox throughout the successful collaborations carried out by their members over the Internet. However, online communities also involve creative arts domains such as animation, video games, and music. Despite their growing popularity, the factors that lead to successful collaborations in these communities are not entirely understood. Fabio Calefato, Giuseppe Iaffaldano, Filippo Lanubile |
GROUP | 1 |
| 2018 | On developers' personality in large-scale distributed projects: the case of the apache ecosystemabstractLarge-scale distributed projects are typically the results of collective efforts performed by multiple developers, each one having a different personality. The study of developers' personalities has the potential of explaining their' behavior in various contexts. For example, the propensity to trust others, a critical factor to the success of global software engineering - has been found to influence positively the result of code reviews in distributed projects. Fabio Calefato, Giuseppe Iaffaldano, Filippo Lanubile, Bogdan Vasilescu |
ICGSE | 1 |
| 2018 | Sentiment polarity detection for software developmentabstractThe role of sentiment analysis is increasingly emerging to study software developers' emotions by mining crowd-generated content within software repositories and information sources. With a few notable exceptions [1][5], empirical software engineering studies have exploited off-the-shelf sentiment analysis tools. However, such tools have been trained on non-technical domains and general-purpose social media, thus resulting in misclassifications of technical jargon and problem reports [2][4]. In particular, Jongeling et al. [2] show how the choice of the sentiment analysis tool may impact the conclusion validity of empirical studies because not only these tools do not agree with human annotation of developers' communication channels, but they also disagree among themselves. Fabio Calefato, Filippo Lanubile, Federico Maiorano, Nicole Novielli |
ICSE | 1 |
| 2018 | A Revision Control System for Image Editing in Collaborative Multimedia DesignabstractRevision control is a vital component in the collaborative development of artifacts such as software code and multimedia. While revision control has been widely deployed for text files, very few attempts to control the versioning of binary files can be found in the literature. This can be inconvenient for graphics applications that use a significant amount of binary data, such as images, videos, meshes, and animations. Existing strategies such as storing whole files for individual revisions or simple binary deltas, respectively consume significant storage and obscure semantic information. To overcome these limitations, in this paper we present a revision control system for digital images that stores revisions in form of graphs. Besides, being integrated with Git, our revision control system also facilitates artistic creation processes in common image editing and digital painting workflows. A preliminary user study demonstrates the usability of the proposed system. Fabio Calefato, Giovanna Castellano, Veronica Rossano |
IV | 1 |
| 2018 | Natural language or not (NLON): a package for software engineering text analysis pipelineabstractThe use of natural language processing (NLP) is gaining popularity in software engineering. In order to correctly perform NLP, we must pre-process the textual information to separate natural language from other information, such as log messages, that are often part of the communication in software engineering. We present a simple approach for classifying whether some textual input is natural language or not. Although our NLoN package relies on only 11 language features and character tri-grams, we are able to achieve an area under the ROC curve performances between 0.976-0.987 on three different data sources, with Lasso regression from Glmnet as our learner and two human raters for providing ground truth. Cross-source prediction performance is lower and has more fluctuation with top ROC performances from 0.913 to 0.980. Compared with prior work, our approach offers similar performance but is considerably more lightweight, making it easier to apply in software engineering text mining pipelines. Our source code and data are provided as an R-package for further improvements. Mika Mäntylä, Fabio Calefato, Maëlick Claes |
MSR | 2 |
| 2018 | A gold standard for emotion annotation in stack overflowabstractSoftware developers experience and share a wide range of emotions throughout a rich ecosystem of communication channels. A recent trend that has emerged in empirical software engineering studies is leveraging sentiment analysis of developers' communication traces. We release a dataset of 4,800 questions, answers, and comments from Stack Overflow, manually annotated for emotions. Our dataset contributes to the building of a shared corpus of annotated resources to support research on emotion awareness in software development. Nicole Novielli, Fabio Calefato, Filippo Lanubile |
MSR | 2 |
| 2018 | Sentiment Polarity Detection for Software Development
Fabio Calefato, Filippo Lanubile, Federico Maiorano, Nicole Novielli |
Empir. Softw. Eng. | 1 |
| 2018 | How to ask for technical help? Evidence-based guidelines for writing questions on Stack Overflow
Fabio Calefato, Filippo Lanubile, Nicole Novielli |
Inf. Softw. Technol. | 1 |
| 2018 | Investigating Crowd Creativity in Online Music CommunitiesabstractCrowd creativity is typically associated with peer-production communities focusing on artistic products like animations, video games, and music, but less frequently to Open Source Software (OSS), despite the fact that also developers must be creative to come up with new solutions to their technical challenges. In this paper, we conduct a study to further the understanding of which factors from prior work in both OSS and art communities are predictive of successful collaboration - defined as reuse of previous songs - in three different songwriting communities, namely Songtree, Splice, and ccMixter. The main findings from this study confirm that the success of collaborations is associated with high community status of recognizable authors and low degree of derivativity of songs. Fabio Calefato, Giuseppe Iaffaldano, Filippo Lanubile, Federico Maiorano |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2017 | A Preliminary Analysis on the Effects of Propensity to Trust in Distributed Software DevelopmentabstractEstablishing trust between developers working atdistant sites facilitates team collaboration in distributed software development. While previous research has focused on how to build and spread trust in absence of direct, face-to-face communication, it has overlooked the effects of the propensity to trust, i.e., the trait of personality representing the individual disposition to perceive the others as trustworthy. In this study, we present a preliminary, quantitative analysis on how the propensity to trust affects the success of collaborations in a distributed project, where thesuccess is represented by pull requests whose code changes and contributions are successfully merged in the project's repository. Fabio Calefato, Filippo Lanubile, Nicole Novielli |
ICGSE | 1 |
| 2016 | The EmoQuest Project: Emotions in Q&A SitesabstractIn this paper, we describe the overall goals and expected contribution of the EmoQuest project. EmoQuest is a three-year multi-disciplinary research project whose main goal is to understand the role of emotions in social media-based knowledge sharing, specifically in online Question and Answer (Q&A) sites. The main research domain of EmoQuest is Computer Supported Cooperative Work (CSCW), with expected outputs in Human-Computer Interaction, Software Engineering, Linguistics, and Psychology. Nicole Novielli, Fabio Calefato, Filippo Lanubile, Giuseppe Mininni, Annarita Taronna |
AVI | 2 |
| 2016 | Moving to Stack Overflow: Best-Answer Prediction in Legacy Developer ForumsabstractContext: Recently, more and more developer communities are abandoning their legacy support forums, moving onto Stack Overflow. The motivations are diverse, yet they typically include achieving faster response time and larger visibility through the access to a modern and very successful infrastructure. One downside of migration, however, is that the history and the crowdsourced knowledge hosted at previous sites remain separated or even get lost if a community decides to abandon completely the legacy developer forum. Fabio Calefato, Filippo Lanubile, Nicole Novielli |
ESEM | 1 |
| 2016 | A Hub-and-Spoke Model for Tool Integration in Distributed DevelopmentabstractToday distributed development depend on an ever-growing plethora of tools that provide a continual stream of updates and place developers into a situation of channel overload and information fragmentation. In this paper, we present our initial work on the definition of a model, named hub-and-spoke, for a loosely-coupled integration of development tools that can help developers cope with these issues, while also increasing their overall situational awareness. Fabio Calefato, Filippo Lanubile |
ICGSE | 1 |
| 2016 | Assessing the impact of real-time machine translation on multilingual meetings in global software projects
Fabio Calefato, Filippo Lanubile, Tayana Conte, Rafael Prikladnicki |
Empir. Softw. Eng. | 1 |
| 2015 | Mining Successful Answers in Stack OverflowabstractRecent research has shown that drivers of success in online question answering encompass presentation quality as well as temporal and social aspects. Yet, we argue that also the emotional style of a technical contribution influences its perceived quality. In this paper, we investigate how Stack Overflow users can increase the chance of getting their answer accepted. We focus on actionable factors that can be acted upon by users when writing an answer and making comments. We found evidence that factors related to information presentation, time and affect all have an impact on the success of answers. Fabio Calefato, Filippo Lanubile, Maria Concetta Marasciulo, Nicole Novielli |
MSR | 1 |
| 2014 | An empirical simulation-based study of real-time speech translation for multilingual global project teamsabstractContext: Real-time speech translation technology is today available but still lacks a complete understanding of how such technology may affect communication in global software projects. Fabio Calefato, Filippo Lanubile, Rafael Prikladnicki, João Henrique Stocker Pinto |
ESEM | 1 |
| 2014 | Mobile Speech Translation for Multilingual Requirements Meetings: A Preliminary StudyabstractCommunication in global software projects usually occurs between native and non-native English speakers with the drawback of an unequal ability to fully understand and contribute to discussions. In this paper, we investigate the adoption of combining speech recognition and machine translation in order to overcome language barriers among stakeholders who are remotely negotiating software requirements. We report our findings from a simulated study where stakeholders communicate speaking three different languages with the help of the Google mobile speech translation service. Fabio Calefato, Filippo Lanubile, Damiano Romita, Rafael Prikladnicki, João Henrique Stocker Pinto |
ICGSE | 1 |
| 2013 | A Preliminary Investigation of the Effect of Social Media on Affective Trust in Customer-Supplier RelationshipsabstractWe present the preliminary results of an ongoing research aimed at investigating the role of social media in the process of trust building, with particular attention to the case of small-medium enterprises (SME). Our findings show that social media contribute to increase the affective trust more than traditional websites. This result suggests that social media have the potential to enhance the business of SMEs other than large companies, by fostering the affective commitment of customers. Fabio Calefato, Filippo Lanubile, Nicole Novielli |
ACII | 1 |
| 2013 | SocialCDE: a social awareness tool for global software teamsabstractWe present SocialCDE, a tool that aims at augmenting Application Lifecycle Management (ALM) platforms with social awareness to facilitate the establishment of interpersonal connections and increase the likelihood of successful interactions by disclosing developers’ personal interests and contextual information. Fabio Calefato, Filippo Lanubile |
ESEC/SIGSOFT FSE | 1 |
| 2012 | Assessing the impact of real-time machine translation on requirements meetings: a replicated experimentabstractOpportunities for global software development are limited in those countries with a lack of English-speaking professionals. Machine translation technology is today available in the form of cross-language web services and can be embedded into multiuser and multilingual chats without disrupting the conversation flow. However, we still lack a thorough understanding of how real-time machine translation may affect communication in global software teams. Fabio Calefato, Filippo Lanubile, Tayana Conte, Rafael Prikladnicki |
ESEM | 1 |
| 2012 | Social Awareness for Global Software TeamsabstractWe hypothesize that information shared on social media can work for distributed software teams as a surrogate of the social awareness, that is information that a person maintains about others in a social or conversational context, gained during informal face-to-face chats. Hence, we have developed a tool that extends a collaborative development environment by aggregating content from social networks and microblogs into developers' workspace. Fabio Calefato, Filippo Lanubile |
ICGSE | 1 |
| 2012 | Computer-mediated communication to support distributed requirements elicitations and negotiations tasks
Fabio Calefato, Daniela E. Damian, Filippo Lanubile |
Empir. Softw. Eng. | 1 |
| 2011 | A Controlled Experiment on the Effects of Machine Translation in Multilingual Requirements MeetingsabstractRequirements engineering is a communication-intensive activity and thus it suffers much from language difficulties in global software projects. Remote requirements meetings can benefit from machine translation as this technology is today available in the form of cross-language chat services. In this paper, we present the design of a controlled experiment to investigate the effects of automatic machine translation services in requirements meetings. Experiment participants, using either Italian or Portuguese as native language, are asked to interact with a communication tool from a distance in order to prioritize and estimate requirements. First results show that real-time machine translation is not disruptive of the conversation flow and is accepted with favor by participants. However, concrete effects are expected to emerge when language barriers are critical. Fabio Calefato, Filippo Lanubile, Rafael Prikladnicki |
ICGSE | 1 |
| 2010 | Investigating the use of tags in collaborative development environments: a replicated studyabstractModern collaborative development environments have recently introduced tagging as a new feature in order to let developers annotate software artifacts with free keywords. Since tagging has the potential to have an impact on task management in software development processes, there is a need to understand how developers use tagging in projects supported by collaborative development environments and how developers' behavior differ from collaborative tagging in the Social Web. Fabio Calefato, Domenico Gendarmi, Filippo Lanubile |
ESEM | 1 |
| 2010 | Can Real-Time Machine Translation Overcome Language Barriers in Distributed Requirements Engineering?abstractIn global software projects work takes place over long distances, meaning that communication will often involve distant cultures with different languages and communication styles that, in turn, exacerbate communication problems. However, being aware of cultural distance is not sufficient to overcome many of the barriers that language differences bring in the way of global project success. In this paper, we investigate the adoption of machine translation (MT) services in synchronous text-based chat in order to overcome any language barrier existing among groups of stakeholders who are remotely negotiating software requirements. We report our findings from a simulated study that compares the efficiency and the effectiveness of two MT services, Google Translate and apertium-service, in translating the messages exchanged during four distributed requirements engineering workshops. The results show that (a) Google Translate produces significantly more adequate translations than Apertium from English to Italian; (b) both services can be used in text-based chat without disrupting real-time interaction. Fabio Calefato, Filippo Lanubile, Pasquale Minervini |
ICGSE | 1 |
| 2009 | Using frameworks to develop a distributed conferencing system: an experience reportabstractAbstract Application frameworks are a powerful means to reduce software development costs while improving quality. However, at the same time they are difficult to select and understand, as well as hard to learn, use, and debug effectively and efficiently. In this paper we report the story of eConference, a distributed conferencing system that was developed as part of a broader research effort. Here we discuss the lessons learned from the evolution of our conferencing tool over four generations, which have been necessary to find good frameworks and build a flexible distributed tool. Copyright © 2009 John Wiley & Sons, Ltd. Fabio Calefato, Filippo Lanubile |
Softw. Pract. Exp. | 1 |
| 2007 | Evolving a text-based conferencing system: An experience reportabstractIn this paper we describe the evolution of eConference, a text-based conferencing system that has turned into a collaborative platform. We draw the lessons learned from the evolution process, as first we changed the underlying communication framework, from the JXTA P2P platform to the XMPP client/server protocol, and then its overall architecture, from traditional plugin to pure-plugin system, built on top of the Eclipse Rich Client Platform. Fabio Calefato, Filippo Lanubile, Mario Scalas |
CollaborateCom | 1 |
| 2007 | A Controlled Experiment on the Effects of Synchronicity in Remote Inspection MeetingsabstractTraditionally, software inspection has largely relied on collocated interaction of inspectors. As companies have begun to turn to distributed software development, meeting in a room has become impractical. In this paper we report on controlled experiment to assess the effect of synchronous and asynchronous communication in remote inspection meetings. Fabio Calefato, Filippo Lanubile, Teresa Mallardo |
ESEM | 1 |
| 2007 | An Empirical Investigation on Text-Based Communication in Distributed Requirements WorkshopsabstractAmong the software development activities, requirements engineering is one of the most communication-intensive and then, its effectiveness is greatly constrained by the geographical distance between stakeholders. For this reason, the need to identify the appropriate task/technology fits to support teams of geographically dispersed stakeholders plays a key role for coping with the lack of physical proximity when developing requirements. In this paper we report on an empirical study that assessed the use of synchronous text-based communication in distributed requirements workshops, as compared to face-to-face (F2F), and the effects of computer-mediated communication (CMC), with respects to the different tasks of distributed requirements elicitation and negotiation. First results show that, in terms of satisfaction with performance, CMC elicitation is a better task/technology fit than CMC negotiation. Furthermore, the general preference for F2F over CMC is due to the strong preference for the F2F negotiation fit over the CMC counterpart. Fabio Calefato, Daniela E. Damian, Filippo Lanubile |
ICGSE | 1 |
| 2004 | Function Clone Detection in Web Applications: A Semiautomated Approach
Fabio Calefato, Filippo Lanubile, Teresa Mallardo |
J. Web Eng. | 1 |