VLDB 2026 Research / reviewers in the wild / expert
Guilherme Avelino 0001
dblp:167/4627 · also Guilherme Amaral Avelino
· DBLP profile ↗
7ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-8203-0638ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Test Co-Evolution in Software Projects: A Large-Scale Empirical StudyabstractABSTRACT The asynchronous evolution of tests and code can compromise software quality and project longevity. To investigate the impact of test and production code co‐evolution, this study analyzes a large‐scale dataset of 526 GitHub repositories written in six programming languages: JavaScript, TypeScript, Java, Python, PHP, and C#. We focus on understanding how tests evolve throughout the software lifecycle and the frequency with which production and test code evolve in sync. By applying clustering algorithms and Pearson's correlation coefficient, we identify different patterns of test co‐evolution between projects. We found a significant correlation between high test co‐evolution and smaller development teams but no significant relationship with the frequency of different maintenance activities (corrective, adaptive, perfective, or multi). Despite this, we identified five distinct test evolution patterns, highlighting diverse approaches to integrating testing practices. This work provides valuable insights into the dynamics of test co‐evolution and its correlation in software maintainability. Charles Miranda, Guilherme Avelino 0001, Pedro de Alcântara dos Santos Neto |
J. Softw. Evol. Process. | 2 |
| 2025 | NoCodeGPT: A No-Code Interface for Building Web Apps With Language ModelsabstractABSTRACT Background Language models are increasingly used by software developers. However, it remains unclear whether their standard chat‐based interfaces are suitable for software development—especially for users with limited programming experience. Objective This work presents a tool, called NoCodeGPT, that provides a customized interface for language models aimed at enabling the implementation of small web applications without writing code. Methods We first conducted an exploratory study in which three participants used ChatGPT to implement a simple web‐based application. After that, we designed and implemented a customized GPT interface, called NoCodeGPT. To evaluate this new interface, we asked 14 students with limited web development experience to build two small web applications using only prompts. Results The exploratory study showed that general‐purpose chat interfaces like ChatGPT are not user‐friendly for application development. One participant, for instance, was unable to complete any proposed user stories. In contrast, results with NoCodeGPT were encouraging: 9 out of 14 participants completed all user stories, while the remaining five completed at least half. Conclusion The standard GPT interface is not well‐suited for novice web developers. In response, we proposed, designed, and implemented a new interface that offers a more accessible experience for building web applications with language models. Mauricio Monteiro, Bruno Castelo Branco, Samuel Silvestre, Guilherme Avelino 0001, Marco Túlio Valente |
Softw. Pract. Exp. | 4 |
| 2024 | Source code expert identification: Models and application
Otávio Cury, Guilherme Avelino 0001, Pedro de Alcântara dos Santos Neto, Marco Túlio Valente, Ricardo Britto 0001 |
Inf. Softw. Technol. | 2 |
| 2022 | Identifying Source Code File ExpertsabstractBackground: In software development, the identification of source code file experts is an important task. Identifying these experts helps to improve software maintenance and evolution activities, such as developing new features, code reviews, and bug fixes. Although some studies have proposed repository-mining techniques to automatically identify source code experts, there are still gaps in this area that can be explored. For example, investigating new variables related to source code knowledge and applying machine learning aiming to improve the performance of techniques to identify source code experts. Aim: The goal of this study is to investigate opportunities to improve the performance of existing techniques to recommend source code files experts. Method: We built an oracle by collecting data from the development history and surveying developers of 113 software projects. Then, we use this oracle to: (i) analyze the correlation between measures extracted from the development history and the developers’ source code knowledge and (ii) investigate the use of machine learning classifiers by evaluating their performance in identifying source code files experts. Results:First Authorship and Recency of Modification are the variables with the highest positive and negative correlations with source code knowledge, respectively. Machine learning classifiers outperformed the linear techniques (F-Measure = 71% to 73%) in the public dataset, but this advantage is not clear in the private dataset, with F-Measure ranging from 55% to 68% for the linear techniques and 58% to 67% for ML techniques. Conclusion: Overall, the linear techniques and the machine learning classifiers achieved similar performance, particularly if we analyze F-Measure. However, machine learning classifiers usually get higher precision while linear techniques obtained the highest recall values. Therefore, the choice of the best technique depends on the user’s tolerance to false positives and false negatives. Otávio Cury, Guilherme Avelino 0001, Pedro de Alcântara dos Santos Neto, Ricardo Britto 0001, Marco Túlio Valente |
ESEM | 2 |
| 2019 | On the abandonment and survival of open source projects: An empirical investigationabstractBackground: Evolution of open source projects frequently depends on a small number of core developers. The loss of such core developers might be detrimental for projects and even threaten their entire continuation. However, it is possible that new core developers assume the project maintenance and allow the project to survive. Aims: The objective of this paper is to provide empirical evidence on: 1) the frequency of project abandonment and survival, 2) the differences between abandoned and surviving projects, and 3) the motivation and difficulties faced when assuming an abandoned project. Method: We adopt a mixed-methods approach to investigate project abandonment and survival. We carefully select 1,932 popular GitHub projects and recover the abandoned and surviving projects, and conduct a survey with developers that have been instrumental in the survival of the projects. Results: We found that 315 projects (16%) were abandoned and 128 of these projects (41%) survived because of new core developers who assumed the project development. The survey indicates that (i) in most cases the new maintainers were aware of the project abandonment risks when they started to contribute; (ii) their own usage of the systems is the main motivation to contribute to such projects; (iii) human and social factors played a key role when making these contributions; and (iv) lack of time and the difficulty to obtain push access to the repositories are the main barriers faced by them. Conclusions: Project abandonment is a reality even in large open source projects and our work enables a better understanding of such risks, as well as highlights ways in avoiding them. Guilherme Avelino 0001, Eleni Constantinou, Marco Túlio Valente, Alexander Serebrenik |
ESEM | 1 |
| 2019 | Measuring and analyzing code authorship in 1 + 118 open source projects
Guilherme Avelino 0001, Leonardo Teixeira Passos, Andre Hora 0001, Marco Túlio Valente |
Sci. Comput. Program. | 1 |
| 2016 | A novel approach for estimating Truck FactorsabstractTruck Factor (TF) is a metric proposed by the agile community as a tool to identify concentration of knowledge in software development environments. It states the minimal number of developers that have to be hit by a truck (or quit) before a project is incapacitated. In other words, TF helps to measure how prepared is a project to deal with developer turnover. Despite its clear relevance, few studies explore this metric. Altogether there is no consensus about how to calculate it, and no supporting evidence backing estimates for systems in the wild. To mitigate both issues, we propose a novel (and automated) approach for estimating TF-values, which we execute against a corpus of 133 popular project in GitHub. We later survey developers as a means to assess the reliability of our results. Among others, we find that the majority of our target systems (65%) have TF ≤ 2. Surveying developers from 67 target systems provides confidence towards our estimates; in 84% of the valid answers we collect, developers agree or partially agree that the TF's authors are the main authors of their systems; in 53% we receive a positive or partially positive answer regarding our estimated truck factors. Guilherme Avelino 0001, Leonardo Teixeira Passos, Andre Hora 0001, Marco Túlio Valente |
ICPC | 1 |