VLDB 2026 Research / reviewers in the wild / expert
João Felipe Pimentel
dblp:161/7127 · also João Felipe N. Pimentel, João Felipe Nicolaci Pimentel
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-6680-7470ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analyzing the adoption of database management systems throughout the history of open source projects
Camila A. Paiva, Raquel Maximino, Frederico Paiva, Rafael Accetta Vieira, Nicole Espanha, João Felipe Pimentel, Igor Scaliante Wiese, Marco Aurélio Gerosa, Igor Steinmacher, Leonardo Murta 0001, Vanessa Braganholo |
Empir. Softw. Eng. | 6 |
| 2024 | How to Support ML End-User Programmers through a Conversational AgentabstractMachine Learning (ML) is increasingly gaining significance for enduser programmer (EUP) applications. However, machine learning end-user programmers (ML-EUPs) without the right background face a daunting learning curve and a heightened risk of mistakes and flaws in their models. In this work, we designed a conversational agent named "Newton" as an expert to support ML-EUPs. Newton's design was shaped by a comprehensive review of existing literature, from which we identified six primary challenges faced by ML-EUPs and five strategies to assist them. To evaluate the efficacy of Newton's design, we conducted a Wizard of Oz within-subjects study with 12 ML-EUPs. Our findings indicate that Newton effectively assisted ML-EUPs, addressing the challenges highlighted in the literature. We also proposed six design guidelines for future conversational agents, which can help other EUP applications and software engineering activities. Emily Judith Arteaga, João Felipe Pimentel, Marco Aurélio Gerosa, Igor Steinmacher, Anita Sarma |
ICSE | 2 |
| 2023 | Tell Me Who Are You Talking to and I Will Tell You What Issues Need Your SkillsabstractSelecting an appropriate task is challenging for newcomers to Open Source Software (OSS) projects. To facilitate task selection, researchers and OSS projects have leveraged machine learning techniques, historical information, and textual analysis to label tasks (a.k.a. issues) with information such as the issue type and domain. These approaches are still far from mainstream adoption, possibly because of a lack of good predictors. Inspired by previous research, we advocate that label prediction might benefit from leveraging metrics derived from communication data and social network analysis (SNA) for issues in which social interaction occurs. Thus, we study how these "social metrics" can improve the automatic labeling of open issues with API domains—categories of APIs used in the source code that solves the issue—which the literature shows that newcomers to the project consider relevant for task selection. We mined data from OSS projects’ repositories and organized it in periods to reflect the seasonality of the contributors’ project participation. We replicated metrics from previous work and added social metrics to the corpus to predict API-domain labels. Social metrics improved the performance of the classifiers compared to using only the issue description text in terms of precision, recall, and F-measure. Precision (0.922) increased by 15.82% and F-measure (0.942) by 15.89% for a project with high social activity. These results indicate that social metrics can help capture the patterns of social interactions in a software project and improve the labeling of issues in an issue tracker. Fabio Santos, Jacob Penney, João Felipe Pimentel, Igor Scaliante Wiese, Igor Steinmacher, Marco Aurélio Gerosa |
MSR | 3 |
| 2023 | Tag that issue: applying API-domain labels in issue tracking systems
Fabio Santos, Joseph Vargovich, Bianca Trinkenreich, Ítalo Santos, Jacob Penney, Ricardo Britto 0001, João Felipe Pimentel, Igor Scaliante Wiese, Igor Steinmacher, Anita Sarma, Marco Aurélio Gerosa |
Empir. Softw. Eng. | 7 |
| 2022 | How to Choose a Task? Mismatches in Perspectives of Newcomers and Existing Contributorsabstract[Background] Selecting an appropriate task is challenging for Open Source Software (OSS) project newcomers and a variety of strategies can help them in this process. [Aims] In this research, we compare the perspective of maintainers, newcomers, and existing contributors about the importance of strategies to support this process. Our goal is to identify possible gulfs of expectations between newcomers who are meant to be helped and contributors who have to put effort into these strategies, which can create friction and impede the usefulness of the strategies. [Method] We interviewed maintainers (n=17) and applied inductive qualitative analysis to derive a model of strategies meant to be adopted by newcomers and communities. Next, we sent a questionnaire (n=64) to maintainers, frequent contributors, and newcomers, asking them to rank these strategies based on their importance. We used the Schulze method to compare the different rankings from the different types of contributors. [Results] Maintainers and contributors diverged in their opinions about the relative importance of various strategies. The results suggest that newcomers want a better contribution process and more support to onboard, while maintainers expect to solve questions using the available communication channels. [Conclusions] The gaps in perspectives between newcomers and existing contributors create a gulf of expectation. OSS communities can leverage our results to prioritize the strategies considered the most important by newcomers. Fabio Santos, Bianca Trinkenreich, João Felipe Pimentel, Igor Scaliante Wiese, Igor Steinmacher, Anita Sarma, Marco Aurélio Gerosa |
ESEM | 3 |
| 2021 | Understanding and improving the quality and reproducibility of Jupyter notebooks
João Felipe Pimentel, Leonardo Murta 0001, Vanessa Braganholo, Juliana Freire |
Empir. Softw. Eng. | 1 |
| 2021 | Recommending Participants for Collaborative Merge SessionsabstractDevelopment of large projects often involves parallel work performed in multiple branches. Eventually, these branches need to be reintegrated through a merge operation. During merge, conflicts may arise and developers need to communicate to reach consensus about the desired resolution. For this reason, including the right developers to a collaborative merge session is fundamental. However, this task can be difficult especially when many different developers have made significant changes on each branch over a large number of files. In this paper, we present TIPMerge, an approach designed to recommend participants for collaborative merge sessions. TIPMerge analyzes the project history and builds a ranked list of developers who are the most appropriate to integrate a pair of branches (Developer Ranking) by considering developers' changes in the branches, in the previous history, and in the dependencies among files across branches. Simply selecting the top developers in such a ranking is easy, but is not effective for collaborative merge sessions as the top developers may have overlapping knowledge. To support collaborative merge, TIPMerge employs optimization techniques to recommend developers with complementary knowledge (Team Recommendation) aiming to maximize joint knowledge coverage. Our results show a mean normalized improvement of 49.5% (median 50.4%) for the joint knowledge coverage with the optimization techniques for assembling teams of three developers for collaborative merge in comparison to choosing the top-3 developers in the ranked list. Catarina de Souza Costa, Jair Figueiredo, João Felipe Pimentel, Anita Sarma, Leonardo Murta 0001 |
IEEE Trans. Software Eng. | 3 |
| 2020 | On the performance of hybrid search strategies for systematic literature reviews in software engineering
Érica Mourão, João Felipe Pimentel, Leonardo Murta 0001, Marcos Kalinowski, Emilia Mendes, Claes Wohlin |
Inf. Softw. Technol. | 2 |
| 2019 | A large-scale study about quality and reproducibility of jupyter notebooksabstractJupyter Notebooks have been widely adopted by many different communities, both in science and industry. They support the creation of literate programming documents that combine code, text, and execution results with visualizations and all sorts of rich media. The self-documenting aspects and the ability to reproduce results have been touted as significant benefits of notebooks. At the same time, there has been growing criticism that the way notebooks are being used leads to unexpected behavior, encourage poor coding practices, and that their results can be hard to reproduce. To understand good and bad practices used in the development of real notebooks, we studied 1.4 million notebooks from GitHub. We present a detailed analysis of their characteristics that impact reproducibility. We also propose a set of best practices that can improve the rate of reproducibility and discuss open challenges that require further research and development. João Felipe Pimentel, Leonardo Murta 0001, Vanessa Braganholo, Juliana Freire |
MSR | 1 |
| 2017 | noWorkflow: a Tool for Collecting, Analyzing, and Managing Provenance from Python ScriptsabstractWe present noWorkflow, an open-source tool that systematically and transparently collects provenance from Python scripts, including data about the script execution and how the script evolves over time. During the demo, we will show how noWorkflow collects and manages provenance, as well as how it supports the analysis of computational experiments. We will also encourage attendees to use noWorkflow for their own scripts. João Felipe Pimentel, Leonardo Murta 0001, Vanessa Braganholo, Juliana Freire |
Proc. VLDB Endow. | 1 |
| 2015 | Software rejuvenation via a multi-agent approach
Heliomar Santos, João Felipe Pimentel, Viviane Torres da Silva, Leonardo Murta 0001 |
J. Syst. Softw. | 2 |