VLDB 2026 Research / reviewers in the wild / expert
Heitor A. X. Costa
dblp:42/7676 · also Heitor Augustus Xavier Costa
· DBLP profile ↗
24ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0002-9903-7414ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 11 since 2021Artificial intelligence and machine learning · 12 · 1 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discovering Patterns in Test Code Refactorings: A Preliminary Study
Railana Santana, Luana Almeida Martins, Larissa Rocha Soares, Carla I. M. Bezerra, Heitor A. X. Costa, Ivan do Carmo Machado |
SEAA (2) | 5 |
| 2025 | Test code refactoring unveiled: where and how does it affect test code quality and effectiveness?
Luana Almeida Martins, Valeria Pontillo, Heitor A. X. Costa, Filomena Ferrucci, Fabio Palomba, Ivan do Carmo Machado |
Empir. Softw. Eng. | 3 |
| 2024 | Tuning Code Smell Prediction Models: A Replication StudyabstractIdentifying code smells in projects is a non-trivial task, and it is often a subjective activity since developers have different understandings about them. The use of machine learning techniques to predict code smells is gaining attention. In this replication study, our goals are: (i) verify if previous model's performance maintain when we extract data from updated systems; and (ii) explore and provide evidences of how the use of different feature engineering and resampling techniques can enhance code smell prediction model's performance. For these purposes, we evaluate four smells: God Class, Refused Bequest, Feature Envy and Long Method. We first replicate a previous study that focus on the algorithm's performance to identify the best models for each smell using a different dataset composed of 30 Java systems. This first experiment provides us a baseline model that is used in the second experiment. In the second experiment, we compare the performance of the baseline model with other models tuned with polynomial features and resample techniques. Our main results are: for datasets with imbalances lower than a ratio of 1:100, such as God Class and Long Method, the use of oversample techniques yielded better results. For datasets with more severe imbalance, like Refused Bequest and Feature Envy, the undersample techniques performed better. The feature selection technique, despite a minor impact on the results, provided insights. For instance, we need new features to represent some code smells, such as Long Method and Feature Envy. Henrique Gomes Nunes, Amanda Santana, Eduardo Figueiredo 0001, Heitor A. X. Costa |
ICPC | 4 |
| 2024 | Predicting merge conflicts considering social and technical assetsabstractAbstract Concurrent contributions to a code base may introduce merge conflicts. Whereas merge conflicts are easy and common to introduce, resolving them is a difficult, time-consuming, and often error-prone task. Previous research concentrated on the emergence of merge conflicts considering technical assets in their analyses and often ignored the social perspective (e.g., developer roles). Our goal is to understand and predict merge conflicts considering social and technical assets. We devise three models for predicting merge conflicts based on common measures used by developers. The first model focuses on the social assets, the second on technical assets, and the third on technical and social assets. To evaluate our predictors, we report on a large-scale empirical study analyzing the histories of 66 real-world software systems. Specifically, we categorize developers into top or occasional contributors at project and merge-scenario level. We found that top contributors at project level and occasional contributors at merge-scenario level cause more merge conflicts than the other roles. Hence, the coordination of top contributors at project level and occasional contributors at merge-scenario level is a good starting point to minimize the occurrence of merge conflicts (especially because when these two developers work on the source branch, the chances of merge conflicts are 32.31%). Overall, we show that predicting merge conflicts incorporating developer roles is possible in practice with high accuracy (0.92) and recall (1.00) when combining technical and social assets, which is vital information to guide improvements on speculative merging techniques. Gustavo Vale, Heitor A. X. Costa, Sven Apel |
Empir. Softw. Eng. | 2 |
| 2024 | An empirical evaluation of RAIDE: A semi-automated approach for test smells detection and refactoring
Railana Santana, Luana Almeida Martins, Tássio Virgínio, Larissa Rocha Soares, Heitor A. X. Costa, Ivan do Carmo Machado |
Sci. Comput. Program. | 5 |
| 2024 | On the diffusion of test smells and their relationship with test code quality of Java projectsabstractAbstract Test smells are considered bad practices that can reduce the test code quality, thus harming software testing goals and maintenance activities. Prior studies have investigated the diffusion of test smells and their impact on test code maintainability. However, we cannot directly compare the outcomes of the studies as most of them use customized datasets. In response, we introduced the TSSM (Test Smells and Structural Metrics) dataset, containing test smells detected using the JNose Test tool and structural metrics (test code and production code) calculated with the CK metrics tool of 13,703 open‐source Java systems from GitHub. In addition, we perform an empirical study to investigate the relationship between test smells and structural metrics of test code and the relationship between test smells on a large‐scale dataset. We split the projects into three clusters to analyze the distribution of test smells, the co‐occurrences among test smells, and the correlation of test smells and structural metrics of test code. The ratio of smelly test classes with a specific test smell is similar among the clusters, but we could observe a significant difference in the number of test smells among them. The test smells Sleepy Test, Mystery Guest, and Resource Optimism rarely occur in the three clusters, and the last two are strongly correlated, indicating that those test smells are more severe than others. Our results point out that most test smells have a moderate correlation with high complexity, large size, and coupling of the test code, indicating that they can also negatively affect its quality. To support further studies, we made our dataset publicly available. Luana Almeida Martins, Heitor A. X. Costa, Ivan do Carmo Machado |
J. Softw. Evol. Process. | 2 |
| 2024 | A comprehensive catalog of refactoring strategies to handle test smells in Java-based systems
Luana Almeida Martins, Taher Ahmed Ghaleb, Heitor A. X. Costa, Ivan do Carmo Machado |
Softw. Qual. J. | 3 |
| 2023 | Hearing the voice of experts: Unveiling Stack Exchange communities' knowledge of test smellsabstractRefactorings are transformations to improve the code design without changing overall functionality and observable behavior. During the refactoring process of smelly test code, practitioners may struggle to identify refactoring candidates and define and apply corrective strategies. This paper reports on an empirical study aimed at understanding how test smells and test refactorings are discussed on the Stack Exchange network. Developers commonly count on Stack Exchange to pick the brains of the wise, i.e., to ‘look up’ how others are completing similar tasks. Therefore, in light of data from the Stack Exchange discussion topics, we could examine how developers understand and perceive test smells, the corrective actions they take to handle them, and the challenges they face when refactoring test code aiming to fix test smells. We observed that developers are interested in others’ perceptions and hands-on experience handling test code issues. Besides, there is a clear indication that developers often ask whether test smells or anti-patterns are either good or bad testing practices than code-based refactoring recommendations. Luana Almeida Martins, Denivan Campos, Railana Santana, Joselito Mota Júnior, Heitor A. X. Costa, Ivan do Carmo Machado |
CHASE | 5 |
| 2023 | Automating Test-Specific Refactoring Mining: A Mixed-Method InvestigationabstractRefactoring is a practice commonly used by developers to restructure the source code without changing its external behavior. Over the last decades, the software engineering research community has been making use of mining software repository techniques to investigate refactoring under multiple perspectives, identifying properties and impact of this practice on source code quality, other than using refactoring data coming from software repositories to build automated recommendation systems. While the current state of the art proposes various automated tools to mine refactoring data, there is still a lack of instruments that may help researchers when mining test-specific refactoring data. The availability of those instruments may enable additional, specialized techniques to support developers while refactoring test code. In this paper, we introduce an approach that extends REFACTORINGMINER-a well-established refactoring mining tool having high precision and recall scores- and is able to detect seven test-specific refactoring operations. We perform mixed-method research to assess capabilities and usefulness of the approach. First, we compare the test-specific refactoring data extracted by the approach against an oracle of 375 test-specific refactorings. Second, we engage with 15 software engineering researchers and apply a technology acceptance model to investigate how they would benefit from our approach. The key results of the study show that our approach reaches 100% and 92.5% of precision and recall scores, respectively. In addition, the approach is considered useful and suitable for various research tasks, including the definition of novel learning models able to recommend test-specific refactoring actions. Luana Almeida Martins, Heitor A. X. Costa, Márcio Ribeiro 0001, Fabio Palomba, Ivan do Carmo Machado |
SCAM | 2 |
| 2021 | A Plugin for Analysis the Usage of Virtual Courses in the Moodle PlatformabstractLearning Management Systems (LMS) consist of a set of virtual tools, focused on the teaching-learning process. In large educational institutions, the number of virtual courses created in this type of environment is significantly high, which raises challenges regarding the administration and maintenance of such courses. One of these challenges is gathering information about the usage of these courses, i.e., whether they are merely used as file repositories, as communication channels with students, or as a framework of learning resources available in the LMS (such as quiz, chat, questionnaires, among others). This paper presents a computational resource (plugin) for the Moodle LMS. This plugin is able to classify, in real time, the type of usage of the virtual courses created in the LMS. In addition, we present the preliminary results obtained with the usage of this plugin at the UFLA (in English, Federal University of Lavras). Alexandre J. C. Silva, Heitor A. X. Costa, Paula Christina Figueira Cardoso, Paulo Afonso Parreira Júnior, Ana Carolina Gondim Inocêncio |
CLEI | 2 |
| 2021 | From Blackboard to the Office: A Look Into How Practitioners Perceive Software Testing EducationabstractThe teaching-learning process may require specific pedagogical approaches to establish a relationship with industry practices. Recently, some studies investigated the educators’ perspectives and the undergraduate courses curriculum to identify potential weaknesses and solutions for the software testing teaching process. However, it is still unclear how the practitioners evaluate the acquisition of knowledge about software testing in undergraduate courses. This study carried out an expert survey with 68 newly graduated practitioners to determine what the industry expects from them and what they learned in academia. The yielded results indicated that those practitioners learned at a similar rate as others with a long industry experience. Also, they studied less than half of the 35 software testing topics collected in the survey and took industry-backed extracurricular courses to complement their learning. Additionally, our findings point out a set of implications for future research, as the respondents’ learning difficulties (e.g., lack of learning sources) and the gap between academic education and industry expectations (e.g., certifications). Luana Almeida Martins, Vinicius Brito, Daniela Soares Feitosa, Larissa Rocha Soares, Heitor A. X. Costa, Ivan do Carmo Machado |
EASE | 5 |
| 2020 | Evolution of quality assessment in SPL: a systematic mappingabstractSoftware product line (SPL) is one of the most recent and effective reuse approaches. SPL derives several products from the core artefacts. SPL engineering includes two processes: domain engineering, which identifies the common and variable features to develop the core artefacts, and application engineering, which reuses the core artefacts to derive products. Once the artefacts are reused across multiple products, quality assessment is necessary to prevent inconsistencies from spreading across all SPL products. There are several frameworks and standards, as ISO/IEC 25010:2011, to evaluate quality characteristics. In this study, the authors provide an overview of the SPL quality assessment. Therefore, they perform a systematic mapping to compile and synthesise data regarding the quality characteristics assessed in studies from 2000 to 2019. The results include the identification of 346 metrics applied in 16 software properties to evaluate three quality characteristics of the ISO/IEC 25010:2011. Additionally, they find the domain engineering evaluation frequently occurs regarding the maintainability characteristic. Moreover, they provide analyses of the: (i) metrics used by programming paradigm, (ii) metrics used by software properties, (iii) software properties evaluated for each quality characteristic, (iv) tools used to extract metrics, (v) systems used as benchmarks, and (vi) datasets used for extracting metrics. Luana Almeida Martins, Paulo Afonso Parreira Júnior, André Pimenta Freire, Heitor A. X. Costa |
IET Softw. | 4 |
| 2019 | Risk Catalogs in Software Project ManagementabstractThe software industry is continuously growing, and projects need to be planned to have a better chance of success. But, planning errors in a project can cause the project to fail. These errors, when there is damage/loss or gain, are called risks and need to be managed. Inadequate risk management can lead to project failure. Therefore, risk management in software design is crucial to its success. In this paper, through research in the literature, catalogs of risks that may occur during the development of software projects are presented. Besides, there are measures defined/identified in the literature to support decision making by project managers, using the GQM method. Valeska Machado, Paulo Afonso Parreira Júnior, Heitor A. X. Costa |
CLEI | 3 |
| 2019 | Challenges and Solutions of Project Management in Distributed Software DevelopmentabstractChallenge of developing software has increased because of several factors, e.g. complexity of these systems and lack of skilled available professional next of development companies. Thus, they should seek alternatives; one of them is global software development. Therefore, software project management should adapt to this reality and resolve challenges not previously found in the “traditional” management. In this paper, we present an initial list of challenges in project management in the context of global software development and solutions proposed by researchers and project managers to try to solve these challenges. We used Literature Systematic Mapping for finding challenges and solutions for software project management. We found 29 papers and used 20 initial papers, totalizing 49 papers. The findings showed 18 challenges which were listed with its solutions. Wallace Moreno, Paulo Afonso Parreira Júnior, Heitor A. X. Costa |
CLEI | 3 |
| 2019 | Extraction of a Software Product Line Using Conditional Compilation - An Exploratory StudyabstractSoftware Product Lines (LPS) is a development approach whose aims is to create a family of software. Despite the increasing interest in software product lines, researches in this area are still very scarce. This hampers broader conclusions about the effective application of principles-based LPS in real systems development. Thus this work describes an experiment involving the extraction of a product line for the TBC-GAAL, educational software developed in Java programming language for teaching Analytic Geometry and Linear Algebra. Using conditional compilation, ten TBC-GAAL features were implemented. The features considered in the experiment were characterized using a set of specific measures for software product lines. Considering the results of this characterization, we highlighted the key challenges involved in extracting features from real software. Patrícia Oliveira, Gustavo Vale, Paulo Afonso Parreira Júnior, Heitor A. X. Costa |
CLEI | 4 |
| 2018 | A systematic mapping study on game-related methods for software engineering education
Maurício R. de A. Souza, Lucas Veado, Renata Teles Moreira, Eduardo Figueiredo 0001, Heitor A. X. Costa |
Inf. Softw. Technol. | 5 |
| 2016 | Heuristic evaluation of the visual accessibility of the moodle Virtual Learning EnvironmentabstractThis paper describes the Heuristic Evaluation performed to investigate the visual accessibility of the Virtual Learning Environment Moodle. The guidelines of the WCAG 2.0 document, related to the visual disability, were used as heuristics. Each heuristic had its own checklist to assist specialists while performing the evaluation. The Heuristic Evaluation was carried out on the pages related to the two main features of Moodle: (i) provide/access files; and (ii) provide/access tasks. As main result, it was noticed that, in general, the Moodle webpages present good accessibility indications. Douglas Costa, Heitor A. X. Costa, Paulo Afonso Parreira Júnior |
CLEI | 2 |
| 2016 | Attributes and metrics of internal quality that impact the external quality of object-oriented software: A systematic literature reviewabstractQuality metrics of software can be categorized into internal quality metrics, external quality metrics, and quality in use metrics. Although existing a close relationship between internal and external quality of software systems, there are no explicit evidences in literature of what are the attributes and metrics of internal quality that impact external quality. Thus, we carried out a systematic literature review for identifying that relationship. After the analysis of 664 papers, 12 papers were studied in depth. As result, we found 65 metrics related primarily to the maintainability, usability, and reliability quality characteristics and the main attributes that impact external metrics are size, coupling, and cohesion. Antônio Maria Pereira de Resende, Paulo Afonso Parreira Júnior, Heitor A. X. Costa |
CLEI | 4 |
| 2015 | Using TDD for developing object-oriented software - A case studyabstractMaintenance of software is accomplished to meet the users' needs of this software (evolution/correction). But, it can become hardest if the source code architecture is difficult to understand. Test Driven Development technique can be used to reduce this difficulty, because it leads the developer to build software with source code simpler. In this paper, this technique is employed to develop software whose functionality is the same of legacy software, but it was developed of way traditional, to obtain more maintainable source code. Software metrics were applied in the source code of legacy and developed software and the results showed improvements in maintainability. Ramon Goncalves, Igor R. Lima, Heitor A. X. Costa |
CLEI | 3 |
| 2015 | Graphical and statistical analysis of the software evolution using coupling and cohesion metrics - An exploratory studyabstractDeveloping software is expensive; thus keep it useful to its users is important. On the other hand, due to constant maintenance performed to meet the changing needs of users, software undergoes degradation of its internal structure, particularly in coupling and cohesion. Monitoring the development of software by using some of its versions can aid Software Engineer with relevant information to guide your maintenance activities. In this paper, we presented a view of the evolution of versions of software. For this, a study was conducted in 10 versions of FindBugs using coupling and cohesion metrics calculated from VizzMaintenance and Metric plug-ins. In this study, we applied the Pearson linear correlation analysis among measurements. The result showed that there is some correlation between these metrics, because coupling metrics directly influenced the cohesion metrics, with undesirable characteristics such as high coupling and low cohesion compromising software quality. Raul Silva, Heitor A. X. Costa |
CLEI | 2 |
| 2014 | Systematic literature review supported by information retrieval techniques: A case studyabstractSystematic Literature Review (SLR) is a means to synthesize relevant and high quality studies related to a specific topic or research questions. In general, a SLR has three phases, and the first one is the Primary Selection in which the selection of studies is usually performed manually reading title, abstract and keywords of each study. The number of published scientific studies has grown, increasing effort in carrying out this review. In this paper, we proposed two strategies to rank studies in decreasing order of importance, for a SLR, regarding the terms in the search string. These strategies are based on the Information Retrieval technique Vector Model. We implemented those strategies and conducted a case study to evaluate their applicability. As results, the second strategy presents 50% of precision on a recall of 80%. Among the contributions of this study, two strategies to rank relevant documents in a SLR, regarding the search string, were proposed and analyzed. Ramon Abílio, Gustavo Vale, Denilson Alves Pereira, Claudiane Oliveira, Flávio Morais, Heitor A. X. Costa |
CLEI | 6 |
| 2013 | Metrics-based Detection of Similar Software (S)
Paloma Oliveira, Hudson Borges, Marco Túlio Valente, Heitor A. X. Costa |
SEKE | 4 |
| 2012 | Mining crosscutting concerns with ComSCId: A rule-based customizable mining toolabstractOne of the first steps when reengineering legacy systems into aspect-oriented ones is to identify the crosscutting concerns (CCC) presented in the architecture of the former; a process known as aspect mining. However, this is a time- consuming and error-prone task when conducted manually. In this paper, we present a customizable mining tool, called ComSCId, which searches for the CCC in legacy Java systems in an automatic way. ComSCId owns a repository which stores all the rules used as base for the mining process. In this repository there are pre-defined rules for some ordinary CCC like persistence, buffering and logging. Moreover, the main characteristic of this repository is its flexibility, since it allows adding new rules or customizing the existing ones to specific contexts or domains. We conducted two studies to evaluate ComSCId and we have observed high percentages of identification coverage when using this tool in an incremental way. Paulo Afonso Parreira Júnior, Wilian Mendes, Valter Vieira de Camargo, Rosângela A. D. Penteado, Heitor A. X. Costa |
CLEI | 5 |
| 2011 | Supporting Software Engineering Education through a Learning Objects and Experience Reports Repository
Rodrigo Pereira dos Santos, Cláudia M. L. Werner, Heitor A. X. Costa, Simone Vasconcelos |
SEKE | 3 |