VLDB 2026 Research / reviewers in the wild / expert
Gustavo Vale
dblp:139/6584 · also Gustavo Andrade Do Vale
· DBLP profile ↗
15ranked-venue papers
6as first author
5since 2021 · last 2024
0000-0002-8879-5797ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 15 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Predicting merge conflicts considering social and technical assetsabstractAbstract Concurrent contributions to a code base may introduce merge conflicts. Whereas merge conflicts are easy and common to introduce, resolving them is a difficult, time-consuming, and often error-prone task. Previous research concentrated on the emergence of merge conflicts considering technical assets in their analyses and often ignored the social perspective (e.g., developer roles). Our goal is to understand and predict merge conflicts considering social and technical assets. We devise three models for predicting merge conflicts based on common measures used by developers. The first model focuses on the social assets, the second on technical assets, and the third on technical and social assets. To evaluate our predictors, we report on a large-scale empirical study analyzing the histories of 66 real-world software systems. Specifically, we categorize developers into top or occasional contributors at project and merge-scenario level. We found that top contributors at project level and occasional contributors at merge-scenario level cause more merge conflicts than the other roles. Hence, the coordination of top contributors at project level and occasional contributors at merge-scenario level is a good starting point to minimize the occurrence of merge conflicts (especially because when these two developers work on the source branch, the chances of merge conflicts are 32.31%). Overall, we show that predicting merge conflicts incorporating developer roles is possible in practice with high accuracy (0.92) and recall (1.00) when combining technical and social assets, which is vital information to guide improvements on speculative merging techniques. Gustavo Vale, Heitor A. X. Costa, Sven Apel |
Empir. Softw. Eng. | 1 |
| 2023 | Yet Another Model! A Study on Model's Similarities for Defect and Code SmellsabstractAbstract Software defect and code smell prediction help developers identify problems in the code and fix them before they degrade the quality or the user experience. The prediction of software defects and code smells is challenging, since it involves many factors inherent to the development process. Many studies propose machine learning models for defects and code smells. However, we have not found studies that explore and compare these machine learning models, nor that focus on the explainability of the models. This analysis allows us to verify which features and quality attributes influence software defects and code smells. Hence, developers can use this information to predict if a class may be faulty or smelly through the evaluation of a few features and quality attributes. In this study, we fill this gap by comparing machine learning models for predicting defects and seven code smells. We trained in a dataset composed of 19,024 classes and 70 software features that range from different quality attributes extracted from 14 Java open-source projects. We then ensemble five machine learning models and employed explainability concepts to explore the redundancies in the models using the top-10 software features and quality attributes that are known to contribute to the defects and code smell predictions. Furthermore, we conclude that although the quality attributes vary among the models, the complexity, documentation, and size are the most relevant. More specifically, Nesting Level Else-If is the only software feature relevant to all models. Geanderson E. dos Santos, Amanda Santana, Gustavo Vale, Eduardo Figueiredo 0001 |
FASE | 3 |
| 2023 | Behind Developer Contributions on Conflicting Merge ScenariosabstractContext: The success of Open Source Software (OSS) projects typically depends on simultaneous contributions of several developers. These contributions often affect the same changing source files and may lead to merge conflicts when integrated. Previous studies investigated the reduction of conflicting merge scenarios. However, empirical evidence on the involvement of OSS contributors in conflicting merge scenarios is scarce. Objective: We aim to fill this gap with a large-scale quantitative study with the goal of understanding: 1) the extent in which OSS contributors are involved in conflicting merge scenarios; 2) characteristics of these contributors; and 3) characteristics of changing source files. Method: We collect both contributor data and contribution data from 66 popular GitHub projects and analyze data of 2972 distinct contributors who were involved in at least one conflicting merge scenario. We rely on both descriptive and inferential statistics to address our research questions. Results: About 80% of the analyzed contributors are involved in only one or two conflicting merge scenarios. Additionally, 42 out of the 66 projects had its top-one contributor as the one mostly involved in conflicting merge scenarios. Finally, only a small set of changing source files are involved in conflicting merge scenarios. Conclusions: We advocate that training the typically small group of contributors involved in conflicting merge scenarios could significantly reduce the number of merge conflicts. Gustavo Vale, Eduardo Fernandes, Eduardo Figueiredo 0001, Sven Apel |
SCAM | 1 |
| 2022 | Challenges of Resolving Merge Conflicts: A Mining and Survey StudyabstractIn collaborative software development, merge conflicts arise when developers integrate concurrent code changes. Practitioners seek to minimize the number of merge conflicts because resolving them is difficult, time consuming, and often an error-prone task. Despite a substantial number of studies investigating merge conflicts, the challenges in merge conflict resolution are not well understood. Our goal is to investigate which factors make merge conflicts longer to resolve in practice. To this end, we performed a two-phase study. First, we analyzed 66 projects containing around 81 thousand merge scenarios, involving 2 million files and over 10 million chunks. For this analysis, we use rank correlation, principal component analysis, multiple regression model, and effect-size analysis to investigate which independent variables (e.g., number of conflicting chunks and files) mostly influence our dependent variable (i.e., time to merge). We found that the number of chunks, lines of code, conflicting chunks, developers involved, conflicting lines of code, conflicting files, and the complexity of the conflicting code influence the merge conflict resolution time. Second, we surveyed 140 developers from our subject projects aiming at cross-validating our results from the first phase of our study. As main results, (i) we found that committing small chunks makes merge conflict resolution faster when leaving other independent variables untouched, (ii) we found evidence that merge scenario characteristics (e.g., the number of lines of code or chunks changed in the merge scenario) are stronger correlated with our dependent variable than merge conflict characteristics (e.g., the number of lines of code or chunks in conflict), (iii) we devise a taxonomy of four types of challenges in merge conflict resolution, and (iv) we observed that the inherent dependencies among conflicting and non-conflicting code is one of the main factors influencing the merge conflict resolution time. Gustavo Vale, Claus Hunsen, Eduardo Figueiredo 0001, Sven Apel |
IEEE Trans. Software Eng. | 1 |
| 2021 | Evaluating T-wise testing strategies in a community-wide dataset of configurable software systems
Fischer Ferreira, Gustavo Vale, João Paulo Diniz, Eduardo Figueiredo 0001 |
J. Syst. Softw. | 2 |
| 2020 | On the relation between Github communication activity and merge conflicts
Gustavo Vale, Angelika Schmid, Alcemir Rodrigues Santos, Eduardo Santana de Almeida, Sven Apel |
Empir. Softw. Eng. | 1 |
| 2019 | Extraction of a Software Product Line Using Conditional Compilation - An Exploratory StudyabstractSoftware Product Lines (LPS) is a development approach whose aims is to create a family of software. Despite the increasing interest in software product lines, researches in this area are still very scarce. This hampers broader conclusions about the effective application of principles-based LPS in real systems development. Thus this work describes an experiment involving the extraction of a product line for the TBC-GAAL, educational software developed in Java programming language for teaching Analytic Geometry and Linear Algebra. Using conditional compilation, ten TBC-GAAL features were implemented. The features considered in the experiment were characterized using a set of specific measures for software product lines. Considering the results of this characterization, we highlighted the key challenges involved in extracting features from real software. Patrícia Oliveira, Gustavo Vale, Paulo Afonso Parreira Júnior, Heitor A. X. Costa |
CLEI | 2 |
| 2019 | On the proposal and evaluation of a benchmark-based threshold derivation method
Gustavo Vale, Eduardo Fernandes, Eduardo Figueiredo 0001 |
Softw. Qual. J. | 1 |
| 2018 | Evaluating domain-specific metric thresholds: an empirical studyabstractSoftware metrics and thresholds provide means to quantify several quality attributes of software systems. Indeed, they have been used in a wide variety of methods and tools for detecting different sorts of technical debts, such as code smells. Unfortunately, these methods and tools do not take into account characteristics of software domains, as the intrinsic complexity of geo-localization and scientific software systems or the simple protocols employed by messaging applications. Instead, they rely on generic thresholds that are derived from heterogeneous systems. Although derivation of reliable thresholds has long been a concern, we still lack empirical evidence about threshold variation across distinct software domains. To tackle this limitation, this paper investigates whether and how thresholds vary across domains by presenting a large-scale study on 3,107 software systems from 15 domains. We analyzed the derivation and distribution of thresholds based on 8 well-known source code metrics. As a result, we observed that software domain and size are relevant factors to be considered when building benchmarks for threshold derivation. Moreover, we also observed that domain-specific metric thresholds are more appropriated than generic ones for code smell detection. Allan Mori, Gustavo Vale, Markos Viggiato, Johnatan Oliveira, Eduardo Figueiredo 0001, Elder Cirilo, Pooyan Jamshidi, Christian Kästner |
TechDebt@ICSE | 2 |
| 2017 | No Code Anomaly is an Island - Anomaly Agglomeration as Sign of Product Line Instabilities
Eduardo Fernandes, Gustavo Vale, Leonardo da Silva Sousa, Eduardo Figueiredo 0001, Alessandro F. Garcia 0001, Jaejoon Lee |
ICSR | 2 |
| 2017 | Identification and Prioritization of Reuse Opportunities with JReuse
Johnatan Oliveira, Eduardo Fernandes, Gustavo Vale, Eduardo Figueiredo 0001 |
ICSR | 3 |
| 2016 | A review-based comparative study of bad smell detection toolsabstractBad smells are symptoms that something may be wrong in the system design or code. There are many bad smells defined in the literature and detecting them is far from trivial. Therefore, several tools have been proposed to automate bad smell detection aiming to improve software maintainability. However, we lack a detailed study for summarizing and comparing the wide range of available tools. In this paper, we first present the findings of a systematic literature review of bad smell detection tools. As results of this review, we found 84 tools; 29 of them available online for download. Altogether, these tools aim to detect 61 bad smells by relying on at least six different detection techniques. They also target different programming languages, such as Java, C, C++, and C#. Following up the systematic review, we present a comparative study of four detection tools with respect to two bad smells: Large Class and Long Method. This study relies on two software systems and three metrics for comparison: agreement, recall, and precision. Our findings support that tools provide redundant detection results for the same bad smell. Based on quantitative and qualitative data, we also discuss relevant usability issues and propose guidelines for developers of detection tools. Eduardo Fernandes, Johnatan Oliveira, Gustavo Vale, Thanis Paiva, Eduardo Figueiredo 0001 |
EASE | 3 |
| 2016 | TDTool: threshold derivation toolabstractSoftware metrics provide basic means to quantify quality of software systems. However, the effectiveness of the measurement process is directly dependent on the definition of reliable thresholds. If thresholds are not properly defined, it is difficult to know, for instance, whether a given metric value indicates a potential problem in a class implementation. There are several methods proposed in literature to derive thresholds for software metrics. However, most of these methods (i) do not respect the skewed distribution of software metrics and (ii) do not provide a supporting tool. Aiming to fill the second gap, we propose a tool, called TDTool, to derive metric thresholds. TDTool is open source and supports four different methods for threshold derivation. This paper presents TDTool architecture and illustrates how to use it. It also presents the thresholds derived using each method based on a benchmark of 33 software product lines. Lucas Veado, Gustavo Vale, Eduardo Fernandes, Eduardo Figueiredo 0001 |
EASE | 2 |
| 2015 | Defining metric thresholds for software product lines: a comparative studyabstractA software product line (SPL) is a set of software systems that share a common and variable set of features. Software metrics provide basic means to quantify several modularity aspects of SPLs. However, the effectiveness of the SPL measurement process is directly dependent on the definition of reliable thresholds. If thresholds are not properly defined, it is difficult to actually know whether a given metric value indicates a potential problem in the feature implementation. There are several methods to derive thresholds for software metrics. However, there is little understanding about their appropriateness for the SPL context. This paper aims at comparing three methods to derive thresholds based on a benchmark of 33 SPLs. We assess to what extent these methods derive appropriate values for four metrics used in product-line engineering. These thresholds were used for guiding the identification of a typical anomaly found in features' implementation, named God Class. We also discuss the lessons learned on using such methods to derive thresholds for SPLs. Gustavo Vale, Danyllo Albuquerque, Eduardo Figueiredo 0001, Alessandro F. Garcia 0001 |
SPLC | 1 |
| 2014 | Systematic literature review supported by information retrieval techniques: A case studyabstractSystematic Literature Review (SLR) is a means to synthesize relevant and high quality studies related to a specific topic or research questions. In general, a SLR has three phases, and the first one is the Primary Selection in which the selection of studies is usually performed manually reading title, abstract and keywords of each study. The number of published scientific studies has grown, increasing effort in carrying out this review. In this paper, we proposed two strategies to rank studies in decreasing order of importance, for a SLR, regarding the terms in the search string. These strategies are based on the Information Retrieval technique Vector Model. We implemented those strategies and conducted a case study to evaluate their applicability. As results, the second strategy presents 50% of precision on a recall of 80%. Among the contributions of this study, two strategies to rank relevant documents in a SLR, regarding the search string, were proposed and analyzed. Ramon Abílio, Gustavo Vale, Denilson Alves Pereira, Claudiane Oliveira, Flávio Morais, Heitor A. X. Costa |
CLEI | 2 |