Everton L. G. Alves

dblp:34/7512 · also Everton L. Galdino, Everton Leandro Galdino Alves · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-2749-0236ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 16 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021
YearPublicationVenuePosition
2026 Beyond Production Code: A Pull Request-Based Study of Technical Debt in Test Code
José Rocha do Amaral Neto, Everton L. G. Alves, Eliane C. Araújo
TechDebt@ICSE2
2025 Zimic: Bridging the Gaps in API Mocking for Typescript Projects
abstract
Mock objects are key for unit and integration testing, as they help isolate tests and prevent flakiness. API mocks, which intercept communications with external services to return simulated responses, must closely align with the actual service to ensure realistic and effective test scenarios. However, for TypeScript projects, there is a notable lack of resources to assist testers in declaring compliant API mocks and identifying incompatibilities early on. Additionally, existing mocking tools offer limited abstractions, making it difficult to address common testing requirements, which in turn leads to increased development time and inconsistent standards. These issues make API mocking a very challenging task, especially for novice testers. In this work, we introduce an approach and tool for API mocks in Typescript projects, Zimic. Our solution enhances mock compliance through static validation, automates type generation from OpenAPI specifications, and provides utilities to support test assertions. We evaluated Zimic in two empirical studies involving students and professional testers. Moreover, we surveyed the participants to get their perspectives on the created API mocks. Our results show that Zimic improves mock compliance by 24.1% compared to existing solutions, while also enhancing completeness and test readability by 27.2% and 15.1%, respectively.
Diego Cruz de Aquino, Everton L. G. Alves
COMPSAC2
2025 Exploring the Use of LLMs to Reduce the Discarding of MBT Test Cases
abstract
Model-Based Testing (MBT) enables the automated generation of test suites from requirement models. However, the frequent changes in agile development often lead teams to indiscriminately discard existing test cases, undermining the efficiency of Test Case Maintenance (TCM). This practice results in the loss of valuable test artifacts and escalates costs due to redundant test generation. Previous research has explored test case reuse through distance functions, but this strategy often suffers from low precision and misclassification. These issues lead to an excessive number of test cases being incorrectly considered reusable. In this paper, we investigate the use of Large Language Models (LLMs) to improve test case management. Through an empirical study on two industrial systems, we analyzed the performance of 13 well-known LLMs in classifying the impact of use case edits using CoT (Chain of Thought)/ToT (Tree of Thought) prompting and Naive-RAG strategies. Our findings indicate that seven of these models effectively reduced the unnecessary discarding of test cases by accurately identifying high-impact requirement changes, achieving a 7% improvement over distance functions. This resulted in a more precise, reliable, and efficient TCM solution within MBT. However, compared to distance-function-based strategies, LLMs exhibited slightly lower recall, performing 6% worse in test case reuse and reduction information loss.
Carlos D. Q. Lima, Everton L. G. Alves, Wilkerson de L. Andrade, Felipe Torres
COMPSAC2
2025 The Hidden Challenges of Merging: A Tool-Based Exploration
abstract
Merging is common in collaborative software development, often leading to conflicts. Code modifications, such as refactorings, may contribute to merge conflicts depending on the approach employed by the merging tools. In this paper, we investigate how code changes over time influence merge conflicts and examine how different merge tools affect their frequency. We analyzed 507,411 merge actions from GitHub Java projects using three distinct merge tools: Git, jFSTMerge, and IntelliMerge. Our findings reveal that nearly 43 percent of Git's conflicting scenarios involved refactorings, and resolving these conflicts required significantly more time. We found that refactorings increase the likelihood of conflicts by roughly 10 times in Git, 12 times in jFSTMerge, and 9 times in IntelliMerge. Additionally, out of 62 refactoring types executed, 33.8 percent were consistently associated with conflicts across all three tools. Regarding the number of developers involved in the history of a conflict, jFSTMerge and IntelliMerge were 28 times and 8 times more likely to result in merge conflicts, respectively. These in-sights may be employed to enhance merging algorithms to better handle these specific types of changes and to guide development teams in mitigating risks by coordinating refactorings, potentially reducing the overall rate of conflicts.
Luciana Q. Leal, Melina Mongiovi, Sabrina Souto, Everton L. G. Alves
SANER4
2025 A qualitative study on refactorings induced by code review
Flávia Coelho, Nikolaos Tsantalis, Tiago Massoni, Everton L. G. Alves
Empir. Softw. Eng.4
2024 A Systematic Literature Review on MBT Test Cases Maintenance
abstract
Model-Based Testing (MBT) can be a valuable tool for software testing, automating test generation from the System Under Test (SUT) models and making the testing process systematic. However, it is common for models to undergo changes during the software lifecycle, which requires adaptive and robust testing strategies. As models evolve, the generated MBT suites require maintenance. While some tests may remain usable, others may become obsolete or require revision. In such scenarios, it is common for parts of an MBT suite to be discarded due to model changes, resulting in additional costs and hindering bug traceability. Therefore, maintaining MBT suites poses a significant challenge to their practical use. This paper presents a Systematic Literature Review (SLR) identifying predominant practices related to MBT suite maintenance, with a focus on strategies for reducing test case discard. The findings reveal that while reuse can prevent the loss of valuable test information, it is fundamentally driven by changes in requirement models, hinging on syntactic differences and the specific constructs of each model's formalism, highlighting the critical role of semantic deepening to minimize test case discard in Test Case Maintenance (TCM) with MBT approaches.
Carlos D. Q. Lima, Everton L. G. Alves, Wilkerson de L. Andrade
COMPSAC2
2024 An Automatic Approach for Uniquely Discovering Actionable Elements for Systematic GUI Testing in Web Applications
abstract
Automated GUI testing can play a crucial role in uncovering faults within web applications. In this context, scriptless testing can streamline the process by automatically generating and executing test cases that explore GUI elements. However, discovering non-repetitive actionable elements within complex web pages remains a challenging task. Often, testers need to employ manual strategies to facilitate element discovery and avoid redundancies. In this paper, we introduce a novel automatic approach, Unique Actionable Element Search (UAES), designed to provide a way to semantically differentiate elements in web applications. By leveraging common HTML properties and semantically expressive locators, UAES streamlines the process of discovering and distinguishing actionable elements, enhancing the efficiency and accuracy of systematic GUI testing. We conducted two empirical studies involving four open-source and twenty industrial web applications. They compared the results of our approach to those obtained using explicit markups included by experienced testers. The results indicate that our approach discovered 94.25% of the elements manually identified by the testers, while also uncovering a significant number of elements (46.86% of all elements discovered) that the testers overlooked.
Thiago Santos de Moura, Francisco Igor de Lima Mendes, Everton L. G. Alves, Ismael Raimundo da Silva Neto, Cláudio de Souza Baptista
QRS3
2023 Assessing and Improving the Quality of Generated Tests in the Context of Maintenance Tasks
abstract
Maintenance tasks often rely on failing test cases, highlighting the importance of well-designed tests for their success. While automatically generated tests can provide higher code coverage and detect faults, it is unclear whether they can be effective in guiding maintenance tasks or if developers fully accept them. In our recent work, we presented the results of a series of empirical studies that evaluated the practical support of generated tests. Our studies with 126 developers showed that automatically generated tests can effectively identify faults during maintenance tasks. Developers were equally effective in creating bug fixes when using manually-written, Evosuite, and Randoop tests. However, developers perceived generated tests as not well-designed and preferred refactored versions of Randoop tests. We plan to enhance Evosuite tests and propose an approach/tool that assesses the quality of generated tests and automatically enhances them. Our research may impact the design and use of generated tests in the context of maintenance tasks.
Wesley B. R. Herculano, Everton L. G. Alves, Melina Mongiovi
COMPSAC2
2021 An Empirical Study on Refactoring-Inducing Pull Requests
abstract
Background: Pull-based development has shaped the practice of Modern Code Review (MCR), in which reviewers can contribute code improvements, such as refactorings, through comments and commits in Pull Requests (PRs). Past MCR studies uniformly treat all PRs, regardless of whether they induce refactoring or not. We define a PR as refactoring-inducing, when refactoring edits are performed after the initial commit(s), as either a result of discussion among reviewers or spontaneous actions carried out by the PR developer. Aims: This mixed study (quantitative and qualitative) explores code reviewing-related aspects intending to characterize refactoring-inducing PRs. Method: We hypothesize that refactoring-inducing PRs have distinct characteristics than non-refactoring-inducing ones and thus deserve special attention and treatment from researchers, practitioners, and tool builders. To investigate our hypothesis, we mined a sample of 1,845 Apache's merged PRs from GitHub, mined refactoring edits in these PRs, and ran a comparative study between refactoring-inducing and non-refactoring-inducing PRs. We also manually examined 2,096 review comments and 1,891 detected refactorings from 228 refactoring-inducing PRs. Results: We found 30.2% of refactoring-inducing PRs in our sample and that they significantly differ from non-refactoring-inducing ones in terms of number of commits, code churn, number of file changes, number of review comments, length of discussion, and time to merge. However, we found no statistical evidence that the number of reviewers is related to refactoring-inducement. Our qualitative analysis revealed that at least one refactoring edit was induced by review in 133 (58.3%) of the refactoring-inducing PRs examined. Conclusions: Our findings suggest directions for researchers, practitioners, and tool builders to improve practices around pull-based code review.
Flávia Coelho, Nikolaos Tsantalis, Tiago Massoni, Everton L. G. Alves
ESEM4
2019 An Empirical Study on the Spreading of Fault Revealing Test Cases in Prioritized Suites
abstract
Code edits are very common during software development. Specially for agile development, these edits need constant validation to avoid functionality regression. In this context, regression test suites are often used. However, regression testing can be very costly. Test case prioritization (TCP) techniques try to reduce this burden by reordering the tests of a given suite aiming at fastening the achievement of a certain testing goal. The literature presents a great number of TCP techniques. Most of the work related to prioritization evaluate the performance of TCP techniques by calculating the rate of test cases that fail per fault (the APFD metric). However, other aspects should be considered when evaluating prioritization results. For instance, the ability to reduce the spreading of failing test cases, since a better grouping often provides more information regarding faults. This paper presents an empirical investigation for evaluating the performance of a set of prioritization techniques comparing APFD and spreading results. Our results show that prioritization techniques generate different APFD and spreading results, being total-statement prioritization the one with the lowest spreading.
Wesley N. M. Torres, Everton L. G. Alves, Patrícia Duarte de Lima Machado
COMPSAC (1)2
2018 Integrating Requirements Specification and Model-Based Testing in Agile Development
abstract
In agile development, Requirements Engineering (RE) and testing have to cope with a number of challenges such as continuous requirement changes and the need for minimal and manageable documentation. In this sense, extensive research has been conducted to automatically generate test cases from (structured) natural language documents using Model-Based Testing (MBT). However, the imposed structure may impair agile practices or test case generation. In this paper, inspired by cooperation with industry partners, we propose CLARET, a notation that allows the creation of use case specifications using natural language to be used as central artifacts for both RE and MBT practices. A tool set supports CLARET specification by checking syntax of use cases structure as well as providing visualization of flows for use case revisions. We also present exploratory studies on the use of CLARET to create RE documents as well as on their use as part of a system testing process based on MBT. Results show that, with CLARET, we can document use cases in a cost-effective way. Moreover, a survey with professional developers shows that CLARET use cases are easy to read and write. Furthermore, CLARET has been successfully applied during specification, development and testing of industrial applications.
Dalton N. Jorge, Patrícia Duarte de Lima Machado, Everton L. G. Alves, Wilkerson de L. Andrade
RE3
2018 Refactoring Inspection Support for Manual Refactoring Edits
abstract
Refactoring is commonly performed manually, supported by regression testing, which serves as a safety net to provide confidence on the edits performed. However, inadequate test suites may prevent developers from initiating or performing refactorings. We propose RefDistiller, a static analysis approach to support the inspection of manual refactorings. It combines two techniques. First, it applies predefined templates to identify potential missed edits during manual refactoring. Second, it leverages an automated refactoring engine to identify extra edits that might be incorrect. RefDistiller also helps determine the root cause of detected anomalies. In our evaluation, RefDistiller identifies 97 percent of seeded anomalies, of which 24 percent are not detected by generated test suites. Compared to running existing regression test suites, it detects 22 times more anomalies, with 94 percent precision on average. In a study with 15 professional developers, the participants inspected problematic refactorings with RefDistiller versus testing only. With RefDistiller, participants located 90 percent of the seeded anomalies, while they located only 13 percent with testing. The results show RefDistiller can help check the correctness of manual refactorings.
Everton L. G. Alves, Myoungkyu Song, Tiago Massoni, Patrícia Duarte de Lima Machado, Miryung Kim
IEEE Trans. Software Eng.1
2017 Test coverage of impacted code elements for detecting refactoring faults: An exploratory study
Everton L. G. Alves, Tiago Massoni, Patrícia Duarte de Lima Machado
J. Syst. Softw.1
2016 Prioritizing test cases for early detection of refactoring faults
abstract
Summary Refactoring edits are error‐prone, requiring cost‐effective testing. Regression test suites are often used as a safety net for decreasing the chances of behavioural changes. Because of the high costs related to handling massive test suites, prioritization techniques can be applied to reorder test case execution, fostering early fault detection. However, traditional prioritization techniques are not specifically designed for detecting refactoring‐related faults. This article proposes refactoring‐based approach (RBA), a refactoring‐aware strategy for prioritizing regression test cases. RBA reorders an existing test sequence, using a set of proposed refactoring fault models that define the refactoring's impact on program methods. Refactoring‐based approach's evaluation shows that it promotes early detection of refactoring faults and outperforms well‐known prioritization techniques in 71% of the cases. Moreover, it prioritizes fault‐revealing test cases close to one another in 73% of the cases, which can be useful for fault localization. Those findings show that RBA can considerably improve prioritization of test cases during perfective evolution, both by increasing fault‐detection rates as well as by helping to pinpoint defects introduced by an incorrect refactoring. Copyright © 2016 John Wiley & Sons, Ltd.
Everton L. G. Alves, Patrícia Duarte de Lima Machado, Tiago Massoni, Miryung Kim
Softw. Test. Verification Reliab.1
2014 RefDistiller: a refactoring aware code review tool for inspecting manual refactoring edits
abstract
Manual refactoring edits are error prone, as refactoring requires developers to coordinate related transformations and understand the complex inter-relationship between affected types, methods, and variables. We present RefDistiller, a refactoring-aware code review tool that can help developers detect potential behavioral changes in manual refactoring edits. It first detects the types and locations of refactoring edits by comparing two program versions. Based on the reconstructed refactoring information, it then detects potential anomalies in refactoring edits using two techniques: (1) a template-based checker for detecting missing edits and (2) a refactoring separator for detecting extra edits that may change a program's behavior. By helping developers be aware of deviations from pure refactoring edits, RefDistiller can help developers have high confidence about the correctness of manual refactoring edits. RefDistiller is available as an Eclipse plug-in at https://sites.google.com/site/refdistiller/ and its demonstration video is available at http://youtu.be/0Iseoc5HRpU.
Everton L. G. Alves, Myoungkyu Song, Miryung Kim
SIGSOFT FSE1
2014 Automatic generation of built-in contract test drivers
Everton L. G. Alves, Patrícia Duarte de Lima Machado, Franklin Ramalho
Softw. Syst. Model.1