VLDB 2026 Research / reviewers in the wild / expert
Baldoino Fonseca dos Santos Neto
dblp:22/7095 · also Baldoino Fonseca
· DBLP profile ↗
36ranked-venue papers
4as first author
16since 2021 · last 2027
0000-0002-0730-0319ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 32 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Foundation models as oracles for refactoring correctness detectionabstractAbstract Refactoring tools in popular Integrated Development Environments (IDEs) can introduce unintended behavioral changes or compilation errors, a persistent challenge that undermines developer trust in automated transformations. Traditional detection approaches rely on handcrafted preconditions, and static and dynamic analyses, yet remain limited in adaptability and can miss subtle correctness issues. This study examines the potential of foundation models to serve as oracles for detecting refactoring bugs in Java programs. We evaluate zero-shot prompting, without task-specific training, across 226 real refactoring bugs collected over more than a decade from widely used Java IDEs ( IntelliJ-IDEA , Eclipse , and NetBeans ), spanning 47 refactoring types. Our results indicate that foundation models can be effective for this task, although performance varies across models. In the first-run setting, GPT-OSS-20B achieved 80.5% accuracy, while GPT-5.4 reached 93.8%. We also evaluated other open-weight and proprietary models: Gemma-4-31B achieved the strongest result among open-weight models, and Gemini-3.1-Pro-Preview achieved the best overall result among all evaluated models. The two primary models were evaluated over five attempts, whereas the additional models were evaluated in a single-run comparison. Metamorphic testing indicates that model predictions remain largely consistent under the tested semantics-preserving perturbations, but these results should be interpreted as robustness evidence rather than as evidence against memorization or data contamination. Beyond detection accuracy, foundation models can provide short explanations that may help support developer inspection, operate across refactoring types without explicitly encoded refactoring-specific rules, and may serve as lightweight triage aids in development workflows. Our findings suggest that foundation models can complement traditional refactoring checks by flagging suspicious transformations for developer inspection. Rohit Gheyi, Rian Melo, Jonhnanthan Oliveira, Márcio Ribeiro 0001, Baldoino Fonseca dos Santos Neto |
Empir. Softw. Eng. | 5 |
| 2025 | Evaluating the Impact of Compression Techniques on the Robustness of CNNs under Natural CorruptionsabstractCompressed deep learning models are crucial for deploying computer vision systems on resource-constrained devices. However, model compression may affect robustness, especially under natural corruption. Therefore, it is important to consider robustness evaluation while validating computer vision systems. This paper presents a comprehensive evaluation of compression techniques—quantization, pruning, and weight clustering—applied individually and in combination to convolutional neural networks (ResNet-50, VGG-19, and MobileNetV2). Using the CIFAR-10-C and CIFAR-100-C datasets, we analyze the trade-offs between robustness, accuracy, and compression ratio. Our results show that certain compression strategies not only preserve but can also improve robustness, particularly on networks with more complex architectures. Utilizing multi-objective assessment, we determine the best configurations, showing that customized technique combinations produce beneficial multi-objective results. This study provides insights into selecting compression methods for robust and efficient deployment of models in corrupted real-world environments. Itallo Patrick Castro Alves Da Silva, Emanuel Adler Medeiros Pereira, Erick A. Barboza, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
ICMLA | 4 |
| 2024 | "Looks Good To Me ;-)": Assessing Sentiment Analysis Tools for Pull Request DiscussionsabstractModern software development relies on cloud-based collaborative platforms (e.g., GitHub and GitLab). In these platforms, developers often employ a pull-based development approach, proposing changes via pull requests and engaging in communication via asynchronous message exchanges. Since communication is key for software development, studies have linked different types of sentiments embedded in the communication to their effects on software projects, such as bug-inducing commits or the non-acceptance of pull requests. In this context, sentiment analysis tools are paramount to detect the sentiment of developers’ messages and prevent potentially harmful impact. Unfortunately, existing state-of-the-art tools vary in terms of the nature of their data collection and labeling processes. Yet, there is no comprehensive study comparing the performance and generalizability of existing tools utilizing a dataset that was designed and systematically curated to this end, and in this specific context. Therefore, in this study, we design a methodology to assess the effectiveness of existing sentiment analysis tools in the context of pull request discussions. For that, we created a dataset that contains ≈ 1.8K manually labeled messages from 36 software projects. The messages were labeled by 19 experts (neuroscientists and software engineers), using a novel and systematic manual classification process designed to reduce subjectivity. By applying these existing tools to the dataset, we observed that while some tools ]perform acceptably, their performance is far from ideal, especially when classifying negative messages. This is interesting since negative sentiment is often related to a critical or unfavorable opinion. We also observed that some messages have characteristics that can make them harder to classify, causing disagreements between the experts and possible misclassifications by the tools, requiring more attention from researchers. Our contributions include valuable resources to pave the way to develop robust and mature sentiment analysis tools that capture/anticipate potential problems during software development. Daniel Coutinho, Luisa Cito, Maria Vitória Lima, Beatriz Arantes, Juliana Alves Pereira, Johny Arriel, João Godinho, Vinicius Martins, Paulo Vítor C. F. Libório, Leonardo Pedrosa Leite, Alessandro F. Garcia 0001, Wesley K. G. Assunção, Igor Steinmacher, Augusto Baffa, Baldoino Fonseca dos Santos Neto |
EASE | 15 |
| 2024 | Enhancing Recommendations of Composite Refactorings based on the PracticeabstractRefactoring is a non-trivial maintenance activity. Developers spend time and effort refactoring code to remove structural problems, i.e., code smells. Recent studies indicated that developers often apply composite refactoring (composite, for short), i.e., two or more interrelated refactorings. However, prior studies revealed that only 10% of composite refactorings are considered complete, i.e., those fully removing code smells. Many incomplete refactorings can even replace or introduce smells, requiring additional effort for their removal later in the project. Moreover, existing refactoring recommendations are not well-detailed and do not alert developers about these possible side effects. To address these gaps, we conducted a large-scale study involving more than 250k refactorings from 42 software projects, including both open-source and closed-source projects. Our goal is to investigate how the most common complete composites are combined and their side effects in the practice. Our results reveal that the current recommendation to apply Extract Method(s) with fine-grained refactoring types needs refinements. We found that certain fine-grained refactorings like Change Variable Types and Change Return Types can introduce up to 45% of Brain Methods when combined with Extract Method(s). Moreover, Ex-tract Method(s) and Move Method(s), a common recommendation to remove Feature Envy, may inadvertently introduce about 30% of Lazy Classes and approximately 70% of Data Classes. Despite these potential side effects, existing refactoring catalogs and tools' recommenders do not alert developers about these side effects. Finally, we consolidate our findings into a catalog to provide clear guidance for developers and researchers on effectively applying composite refactorings to fully remove code smells. Ana Carla Bibiano, Daniel Coutinho, Anderson G. Uchôa, Wesley K. G. Assunção, Alessandro F. Garcia 0001, Rafael Maiani de Mello, Thelma Elita Colanzi, Daniel Oliveira 0005, Audrey Vasconcelos, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
SCAM | 10 |
| 2024 | Towards effective gamification of existing systems: method and experience report
Anderson G. Uchôa, Rafael Maiani de Mello, Jairo Souza, Leopoldo Teixeira, Baldoino Fonseca dos Santos Neto, Alessandro F. Garcia 0001 |
Softw. Qual. J. | 5 |
| 2023 | Beyond the Code: Investigating the Effects of Pull Request Conversations on Design DecayabstractBackground: Code development is done collaboratively in platforms such as GitHub and GitLab, following a pull-based development model. In this model, developers actively communicate and share their knowledge through conversations. Pull request conversations are affected by social aspects such as communication dynamics among developers, discussion content, and organizational dynamics. Despite prior studies indicating that social aspects indeed impact software quality, it is still unknown to what extent social aspects influence design decay during software development. Thus, since social aspects are intertwined with design and implementation decisions, there is a need for investigating how social aspects contribute to avoiding, reducing, or accelerating design decay. Aims: To fill this gap, we performed a study aimed at investigating the effects of pull request conversation on design decay. Method: We investigated 10,746 pull request conversations from 11 open-source systems, characterizing in terms of three different social aspects: discussion content, organizational and communication dynamics. We considered 18 social metrics to these three social aspects, and analyzed how they associate with design decay. We used a statistical approach to assess which social metrics are able to discriminate between impactful and unimpactful pull requests. Then, we employed a multiple logistic regression model to evaluate the influence of each social metric per social aspect in the presence of each other on design decay. Finally, we also observed how the combination of all social metrics influences the design decay. Results: Our findings reveal that social metrics related to the size and duration of a discussion, the presence of design-related keywords, the team size, and gender diversity can be used to discriminate between design impactful and unimpactful pull requests. Organizational growth and gender diversity prevent decay. Each software community has its unique aspects that can be used to detect and prevent design decay. Also, design improvements can be accomplished by timely feedback, engaged communication, and design-oriented discussions with the contribution of multiple participants who provide significant comments. Conclusion: The social aspects related to pull request conversations are useful indicators of design decay. Caio Barbosa, Anderson G. Uchôa, Daniel Coutinho, Wesley K. G. Assunção, Anderson Oliveira, Alessandro F. Garcia 0001, Baldoino Fonseca dos Santos Neto, Matheus Rabelo, José Eric Coelho, Eryka Carvalho, Henrique Santos 0003 |
ESEM | 7 |
| 2023 | Manual Tests Do Smell! Cataloging and Identifying Natural Language Test SmellsabstractBackground: Test smells indicate potential problems in the design and implementation of automated software tests that may negatively impact test code maintainability, coverage, and reliability. When poorly described, manual tests written in natural language may suffer from related problems, which enable their analysis from the point of view of test smells. Despite the possible prejudice to manually tested software products, little is known about test smells in manual tests, which results in many open questions regarding their types, frequency, and harm to tests written in natural language. Aims: Therefore, this study aims to contribute to a catalog of test smells for manual tests. Method: We perform a two-fold empirical strategy. First, an exploratory study in manual tests of three systems: the Ubuntu Operational System, the Brazilian Electronic Voting Machine, and the User Interface of a large smartphone manufacturer. We use our findings to propose a catalog of eight test smells and identification rules based on syntactical and morphological text analysis, validating our catalog with 24 in-company test engineers. Second, using our proposals, we create a tool based on Natural Language Processing (NLP) to analyze the subject systems' tests, validating the results. Results: We observed the occurrence of eight test smells. A survey of 24 in-company test professionals showed that 80.7% agreed with our catalog definitions and examples. Our NLP-based tool achieved a precision of 92%, recall of 95%, and f-measure of 93.5%, and its execution evidenced 13,169 occurrences of our cataloged test smells in the analyzed systems. Conclusion: We contribute with a catalog of natural language test smells and novel detection strategies that better explore the capabilities of current NLP mechanisms with promising results and reduced effort to analyze tests written in different idioms. Elvys Soares, Manoel Aranda III, Naelson Oliveira, Márcio Ribeiro 0001, Rohit Gheyi, Emerson Souza, Ivan do Carmo Machado, André L. M. Santos, Baldoino Fonseca dos Santos Neto, Rodrigo Bonifácio |
ESEM | 9 |
| 2023 | The untold story of code refactoring customizations in practiceabstractRefactoring is a common software maintenance practice. The literature defines standard code modifications for each refactoring type and popular IDEs provide refactoring tools aiming to support these standard modifications. However, previous studies indicated that developers either frequently avoid using these tools or end up modifying and even reversing the code automatically refactored by IDEs. Thus, developers are forced to manually apply refactorings, which is cumbersome and error-prone. This means that refactoring support may not be entirely aligned with practical needs. The improvement of tooling support for refactoring in practice requires understanding in what ways developers tailor refactoring modifications. To address this issue, we conduct an analysis of 1,162 refactorings composed of more than 100k program modifications from 13 software projects. The results reveal that developers recurrently apply patterns of additional modifications along with the standard ones, from here on called patterns of customized refactorings. For instance, we found customized refactorings in 80.77% of the Move Method instances observed in the software projects. We also investigated the features of refactoring tools in popular IDEs and observed that most of the customization patterns are not fully supported by them. Additionally, to understand the relevance of these customizations, we conducted a survey with 40 developers about the most frequent customization patterns we found. Developers confirm the relevance of customization patterns and agree that improvements in IDE's refactoring support are needed. These observations highlight that refactoring guidelines must be updated to reflect typical refactoring customizations. Also, IDE builders can use our results as a basis to enable a more flexible application of automated refactorings. For example, developers should be able to choose which method must handle exceptions when extracting an exception code into a new method. Daniel Oliveira 0005, Wesley K. G. Assunção, Alessandro F. Garcia 0001, Ana Carla Bibiano, Márcio Ribeiro 0001, Rohit Gheyi, Baldoino Fonseca dos Santos Neto |
ICSE | 7 |
| 2023 | Seeing confusion through a new lens: on the impact of atoms of confusion on novices' code comprehension
José Aldo Silva da Costa, Rohit Gheyi, Fernando Castor Filho, Pablo Roberto Fernandes de Oliveira, Márcio Ribeiro 0001, Baldoino Fonseca dos Santos Neto |
Empir. Softw. Eng. | 6 |
| 2023 | Towards a better understanding of the mechanics of refactoring detection tools
Jonhnanthan Oliveira, Rohit Gheyi, Leopoldo Teixeira, Márcio Ribeiro 0001, Osmar Leandro, Baldoino Fonseca dos Santos Neto |
Inf. Softw. Technol. | 6 |
| 2022 | Lint-Based Warnings in Python Code: Frequency, Awareness and RefactoringabstractPython is a popular programming language characterized by its simple syntax and easy learning curve. Like many languages, Python has a set of best practices that should be followed to avoid bugs and improve other quality attributes (such as maintenance and readability). In this context, non-compliance to these practices can be detected by using linting tools. Previous work conducted studies to better understand the frequency of a class of problems that can be found using Python linters: warnings, here named as lint-based warnings. However, they either rely on small datasets or focus on few domains, such as machine learning or web-systems projects. In this paper, we provide a mixed-method study where we analyze the frequency of six lint-based warnings in 1,119 different open-source general-purpose Python projects. To go further, we also conduct a survey to check whether developers are aware of the lint-based warnings we study here. In particular, we intend to check whether they are able to identify the six lint-based warnings. To remove the lint-based warnings, we suggest the application of simple refactorings. Last but not least, we evaluate the suggestions by submitting pull requests to remove lint-based warnings from open-source projects. Our results show that 39% of the 1,119 projects have at least one lint-based warning. After analyzing the survey data, we also show that developers prefer Python code without lint-based warnings. Regarding the pull requests, we achieve a 71.8% of acceptance rate. Naelson Oliveira, Márcio Ribeiro 0001, Rodrigo Bonifácio, Rohit Gheyi, Igor Scaliante Wiese, Baldoino Fonseca dos Santos Neto |
SCAM | 6 |
| 2022 | Developers' perception matters: machine learning to detect developer-sensitive smells
Daniel Oliveira 0005, Wesley K. G. Assunção, Alessandro F. Garcia 0001, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
Empir. Softw. Eng. | 4 |
| 2022 | Developers' viewpoints to avoid bug-introducing changes
Jairo Souza, Rodrigo Lima 0002, Baldoino Fonseca dos Santos Neto, Bruno Cartaxo, Márcio Ribeiro 0001, Gustavo Pinto 0001, Rohit Gheyi, Alessandro F. Garcia 0001 |
Inf. Softw. Technol. | 3 |
| 2021 | Look Ahead! Revealing Complete Composite Refactorings and their Smelliness EffectsabstractRecent studies have revealed that developers often apply composite refactorings (or, simply, composites). A composite consists of two or more interrelated refactorings applied together. Previous studies investigated the effect of composites on code smells. A composite is considered “complete” whenever it completely removes one target code smell. They proposed descriptions of complete composites with recommendations to remove certain code smell types, such as Long Methods and Feature Envies. These studies also present different recommendations to remove the same code smell type. However, these studies: (i) are limited to composites only consisting of a small subset of Fowler's refactoring types, (ii) do not detail the scenarios in which each recommendation can be applied to remove the code smell, and (iii) fail in reporting possible side effects of the described composites, such as adversely introducing certain smell types. This paper aims to cover these limitations by performing a systematic analysis of 618 complete composites on removing four common smell types identified in 20 software projects. Our results indicated that: (i) 64% complete composites consisted of refactoring types not covered by existing descriptions of complete composites, and (ii) 36% complete composites formed by Extract Methods can introduce Feature Envies and Intensive Couplings. This information is not documented by existing descriptions, and it can alert developers about alternatives to remove Feature Envy, mainly in methods that are fully envious. These results suggest existing descriptions of complete composites should be either revisited or enhanced to explicitly highlight known side effects. We present a catalog of composites with details about side effects, recommendations to remove or minimize them, and some scenarios in which each recommendation can be applied to remove the code smell. Our catalog can be useful to improve existing tooling support for refactorings, such as IDEs, informing about possible side effects when refactorings are composed. Ana Carla Bibiano, Wesley K. G. Assunção, Daniel Coutinho, Kleber Santos, Vinícius Soares, Rohit Gheyi, Alessandro F. Garcia 0001, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Daniel Oliveira 0005, Caio Barbosa, João Lucas Marques, Anderson Oliveira |
ICSME | 8 |
| 2021 | Evaluating refactorings for disciplining #ifdef annotations: An eye tracking study with novices
José Aldo Silva da Costa, Rohit Gheyi, Márcio Ribeiro 0001, Sven Apel, Vander Alves, Baldoino Fonseca dos Santos Neto, Flávio Medeiros, Alessandro F. Garcia 0001 |
Empir. Softw. Eng. | 6 |
| 2021 | Identifying method-level mutation subsumption relations using Z3
Rohit Gheyi, Márcio Ribeiro 0001, Beatriz Souza, Marcio Augusto Guimarães, Leonardo Fernandes, Marcelo d'Amorim, Vander Alves, Leopoldo Teixeira, Baldoino Fonseca dos Santos Neto |
Inf. Softw. Technol. | 9 |
| 2020 | Is Exceptional Behavior Testing an Exception?: An Empirical Assessment Using Java Automated TestsabstractSoftware testing is a crucial activity to check the internal quality of a software. During testing, developers often create tests for the normal behavior of a particular functionality (e.g., was this file properly uploaded to the cloud?). However, little is known whether developers also create tests for the exceptional behavior (e.g., what happens if the network fails during the file upload?). To minimize this knowledge gap, in this paper we design and perform a mixed-method study to understand how 417 open source Java projects are testing the exceptional behavior using the JUnit and TestNG frameworks, and the AssertJ library. We found that 254 (60.91%) projects have at least one test method dedicated to test the exceptional behavior. We also found that the number of test methods for exceptional behavior with respect to the total number of test methods lies between 0% and 10% in 317 (76.02%) projects. Also, 239 (57.31%) projects test only up to 10% of the used exceptions in the System Under Test (SUT). When it comes to mobile apps, we found that, in general, developers pay less attention to exceptional behavior tests when compared to desktop/server and multi-platform developers. In general, we found more test methods covering custom exceptions (the ones created in the own project) when compared to standard exceptions available in the Java Development Kit (JDK) or in third-party libraries. To triangulate the results, we conduct a survey with 66 developers from the projects we study. In general, the survey results confirm our findings. In particular, the majority of the respondents agrees that developers often neglect exceptional behavior tests. As implications, our numbers might be important to alert developers that more effort should be placed on creating tests for the exceptional behavior. Francisco Dalton, Márcio Ribeiro 0001, Gustavo Pinto 0001, Leonardo Fernandes, Rohit Gheyi, Baldoino Fonseca dos Santos Neto |
EASE | 6 |
| 2020 | On the Performance and Adoption of Search-Based Microservice Identification with toMicroservicesabstractThe expensive maintenance of legacy systems leads companies to migrate such systems to microservice architectures. This migration requires the identification of system's legacy parts to become microservices. However, the successful identification of microservices, which are promising to be adoptable in practice, requires the simultaneous satisfaction of many criteria, such as coupling, cohesion, reuse and communication overhead. Search-based microservice identification has been recently investigated to address this problem. However, state-of-the-art search-based approaches are limited as they only consider one or two criteria (namely cohesion and coupling), possibly not fulfilling the practical needs of developers. To overcome these limitations, we propose toMicroservices, a many-objective search-based approach that considers five criteria, the most cited by practitioners in recent studies. Our approach was evaluated in a real-life industrial legacy system undergoing a microservice migration process. The performance of toMicroservices was quantitatively compared to a baseline. We also gathered qualitative evidence based on developers' perceptions, who judged the adoptability of the recommended microservices. The results show that our approach is both: (i) very similar to the most recent proposed approach on optimizing the traditional criteria of coupling and cohesion, but (ii) much better when taking into account all the five criteria. Finally, most of the microservice candidates were considered adoptable by practitioners. Alessandro F. Garcia 0001, Thelma Elita Colanzi, Wesley K. G. Assunção, Juliana Alves Pereira, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Maria Julia de Lima, Carlos José Pereira de Lucena |
ICSME | 6 |
| 2020 | How Does Incomplete Composite Refactoring Affect Internal Quality Attributes?abstractProgram refactoring consists of code changes applied to improve the internal structure of a program and, as a consequence, its comprehensibility. Recent studies indicate that developers often perform composite refactorings, i.e., a set of two or more interrelated single refactorings. Recent studies also recommend certain patterns of composite refactorings to fully remove poor code structures, i.e, code smells, thus further improving the program comprehension. However, other recent studies report that composite refactorings often fail to fully remove code smells. Given their failure to achieve this purpose, these composite refactorings are considered incomplete, i.e, they are not able to entirely remove a smelly structure. Unfortunately, there is no study providing an in-depth analysis of the incompleteness nature of many composites and their possibly partial impact on improving, maybe decreasing, internal quality attributes. This paper identifies the most common forms of incomplete composites, and their effect on quality attributes, such as coupling and cohesion, which are known to have an impact on program comprehension. We analyzed 353 incomplete composite refactorings in 5 software projects, two common code smells (Feature Envy and God Class), and four internal quality attributes. Our results reveal that incomplete composite refactorings with at least one Extract Method are often (71%) applied without Move Methods on smelly classes. We have also found that most incomplete composite refactorings (58%) tended to at least maintain the internal structural quality of smelly classes, thereby not causing more harm to program comprehension. We also discuss the implications of our findings to the research and practice of composite refactoring. Ana Carla Bibiano, Vinícius Soares, Daniel Coutinho, Eduardo Fernandes, João Lucas Correia, Kleber Santos, Anderson Oliveira, Alessandro F. Garcia 0001, Rohit Gheyi, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Caio Barbosa, Daniel Oliveira 0005 |
ICPC | 10 |
| 2020 | Refactoring from 9 to 5? What and When Employees and Volunteers Contribute to OSSabstractIn this paper we characterize the contributions made by employees (developers that work for GitHub, the company) and volunteers (developers that use GitHub, the platform) to OSS projects maintained by GitHub (the company) on GitHub (the platform). By mining activities performed in five well-known company-owned OSS projects, we investigate what they do and when they do it. We found that the majority of the volunteers' contributions are related to reengineering (e.g., refactoring), while employees focus more on management (e.g., documentation). When it comes to the working hours, we found that contributions are made mostly from 9am-5pm, even for the volunteers. Luiz Felipe Dias, Caio Barbosa, Gustavo Pinto 0001, Igor Steinmacher, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Christoph Treude, Daniel Alencar da Costa |
VL/HCC | 5 |
| 2020 | On Relating Technical, Social Factors, and the Introduction of BugsabstractAs collaborative coding environments make it easier to contribute to software projects, the number of developers involved in these projects keeps increasing. This increase makes it more difficult for code reviewers to deal with buggy contributions. Collaborative environments like GitHub provide a rich source of data on developers' contributions. Such data can be used to extract information about developers regarding technical (e.g., their experience) and social (e.g., their interactions) factors. Recent studies analyzed the influence of these factors on different activities of software development. However, there is still room for improvement on the relation between these factors and the introduction of bugs. We present a broader study, including 8 projects from different domains and 6,537 bug reports, on relating five technical, three social factors, and the introduction of bugs. The results indicate that technical and social factors can discriminate between buggy and clean commits. But, the technical factors are more determining than social ones. Particularly, the developers' habits of not following technical contribution norms and the developer's commit bugginess are associated with an increase on commit bugginess. On the other hand, project's establishment, ownership level of developers' commit, and social influence are related to a lower chance of introducing bugs. Filipe Falcão, Caio Barbosa, Baldoino Fonseca dos Santos Neto, Alessandro F. Garcia 0001, Márcio Ribeiro 0001, Rohit Gheyi |
SANER | 3 |
| 2020 | Mutating code annotations: An empirical evaluation on Java and C# programs
Pedro Pinheiro, José Carlos Viana, Márcio Ribeiro 0001, Leonardo Fernandes, Fabiano Cutigi Ferrari, Rohit Gheyi, Baldoino Fonseca dos Santos Neto |
Sci. Comput. Program. | 7 |
| 2019 | A Quantitative Study on Characteristics and Effect of Batch Refactoring on Code SmellsabstractBackground: Code refactoring aims to improve code structures via code transformations. A single transformation rarely suffices to fully remove code smells that reveal poor code structures. Most transformations are applied in batches, i.e. sets of interrelated transformations, rather than in isolation. Nevertheless, empirical knowledge on batch application, or batch refactoring, is scarce. Such scarceness helps little to improve current refactoring practices. Aims: We analyzed 57 open and closed software projects. We aimed to understand batch application from two perspectives: characteristics that typically constitute a batch (e.g., the variety of transformation types employed), and the batch effect on smells. Method: We analyzed 19 smell types and 13 transformation types. We identified 4,607 batches, each applied by the same developer on the same code element (method or class); we expected to have batches whose transformations are closely interrelated. We computed (1) the frequency in which five batch characteristic manifest, (2) the probability of each batch characteristics to remove smells, and (3) the frequency in which batches introduce and remove smells. Results: Most batches are quite simple: although most batches are applied on more than one method (90%), they are usually composed of the same transformation type (72%) and only two transformations (57%). Batches applied on a single method are 2.6 times more prone to fully remove smells than batches affecting more than one method. Surprisingly, batches mostly ended up introducing (51%) or not fully removing (38%) smells. Conclusions: The batch simplicity suggests that developers have sub-explored the combinations of transformations within a batch. We summarized some batches that may fully remove smells, so that developers can incorporate them into current refactoring practices. Ana Carla Bibiano, Eduardo Fernandes, Daniel Oliveira 0005, Alessandro F. Garcia 0001, Marcos Kalinowski, Baldoino Fonseca dos Santos Neto, Roberto Oliveira 0003, Anderson Oliveira, Diego Cedrim |
ESEM | 6 |
| 2019 | Software Engineering Research Community Viewpoints on Rapid ReviewsabstractBackground: One of the most important current challenges of Software Engineering (SE) research is to provide relevant evidence to practice. In health related fields, Rapid Reviews (RRs) have shown to be an effective method to achieve that goal. However, little is known about how the SE research community perceives the potential applicability of RRs. Aims: The goal of this study is to understand the SE research community viewpoints towards the use of RRs as a means to provide evidence to practitioners. Method: To understand their viewpoints, we invited 37 researchers to analyze 50 opinion statements about RRs, and rate them according to what extent they agree with each statement. Q-Methodology was employed to identify the most salient viewpoints, represented by the so called factors. Results: Four factors were identified: Factor A groups undecided researchers that need more evidence before using RRs; Researchers grouped in Factor B are generally positive about RRs, but highlight the need to define minimum standards; Factor C researchers are more skeptical and reinforce the importance of high quality evidence; Researchers aligned to Factor D have a pragmatic point of view, considering RRs can be applied based on the context and constraints faced by practitioners. Conclusions: In conclusion, although there are opposing viewpoints, there are also some common grounds. For example, all viewpoints agree that both RRs and Systematic Reviews can be poorly or well conducted. Bruno Cartaxo, Gustavo Pinto 0001, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Pedro Pinheiro, Maria Teresa Baldassarre, Sérgio Soares |
ESEM | 3 |
| 2019 | Do Research and Practice of Code Smell Identification Walk Together? A Social Representations AnalysisabstractContext: It is frequently claimed the need for bridging the gap between software engineering research and practice. In this sense, the theory of social representations may be useful to characterize the actual concerns of software developers. It comprises the system of values, behaviors, and practices of communities regarding a particular social object, such as the task of smell identification. Aim: To characterize the social representations of smell identification by software developers. Method: Based on the answers given to a question-naire, we analyzed the associations made by the developers about smell identification, i.e., what immediately comes to their minds when they think about this task. Results: We found that developers strongly associate smell identification with the practice of smell removal and with the incidence of bugs. They also frequently associate the task with the practice of inspection and with the need of having individual skills. Besides, we verified that the current state of the art on smell identification partially address the social representations of the software developers. Conclusion: There is a considerable gap between the research of smell identification and its practice. We propose directions to mitigating this gap. Rafael Maiani de Mello, Anderson G. Uchôa, Roberto Oliveira 0003, Willian Nalepa Oizumi, Jairo Souza, Kleyson Mendes, Daniel Oliveira 0005, Baldoino Fonseca dos Santos Neto, Alessandro F. Garcia 0001 |
ESEM | 8 |
| 2018 | Identifying design problems in the source code: a grounded theoryabstractThe prevalence of design problems may cause re-engineering or even discontinuation of the system. Due to missing, informal or outdated design documentation, developers often have to rely on the source code to identify design problems. Therefore, developers have to analyze different symptoms that manifest in several code elements, which may quickly turn into a complex task. Although researchers have been investigating techniques to help developers in identifying design problems, there is little knowledge on how developers actually proceed to identify design problems. In order to tackle this problem, we conducted a multi-trial industrial experiment with professionals from 5 software companies to build a grounded theory. The resulting theory offers explanations on how developers identify design problems in practice. For instance, it reveals the characteristics of symptoms that developers consider helpful. Moreover, developers often combine different types of symptoms to identify a single design problem. This knowledge serves as a basis to further understand the phenomena and advance towards more effective identification techniques. Leonardo da Silva Sousa, Anderson Oliveira, Willian Nalepa Oizumi, Simone D. J. Barbosa, Alessandro F. Garcia 0001, Jaejoon Lee, Marcos Kalinowski, Rafael Maiani de Mello, Baldoino Fonseca dos Santos Neto, Roberto Oliveira 0003, Carlos José Pereira de Lucena, Rodrigo B. de Paes |
ICSE | 9 |
| 2018 | Are you smelling it? Investigating how similar developers detect code smells
Mario Hozano, Alessandro F. Garcia 0001, Baldoino Fonseca dos Santos Neto, Evandro de Barros Costa |
Inf. Softw. Technol. | 3 |
| 2018 | Discipline Matters: Refactoring of Preprocessor Directives in the #ifdef HellabstractThe C preprocessor is used in many C projects to support variability and portability. However, researchers and practitioners criticize the C preprocessor because of its negative effect on code understanding and maintainability and its error proneness. More importantly, the use of the preprocessor hinders the development of tool support that is standard in other languages, such as automated refactoring. Developers aggravate these problems when using the preprocessor in undisciplined ways (e.g., conditional blocks that do not align with the syntactic structure of the code). In this article, we proposed a catalogue of refactorings and we evaluated the number of application possibilities of the refactorings in practice, the opinion of developers about the usefulness of the refactorings, and whether the refactorings preserve behavior. Overall, we found 5,670 application possibilities for the refactorings in 63 real-world C projects. In addition, we performed an online survey among 246 developers, and we submitted 28 patches to convert undisciplined directives into disciplined ones. According to our results, 63 percent of developers prefer to use the refactored (i.e., disciplined) version of the code instead of the original code with undisciplined preprocessor usage. To verify that the refactorings are indeed behavior preserving, we applied them to more than 36 thousand programs generated automatically using a model of a subset of the C language, running the same test cases in the original and refactored programs. Furthermore, we applied the refactorings to three real-world projects: BusyBox, OpenSSL, and SQLite. This way, we detected and fixed a few behavioral changes, 62 percent caused by unspecified behavior in the C programming language. Flávio Medeiros, Márcio Ribeiro 0001, Rohit Gheyi, Sven Apel, Christian Kästner, Bruno Ferreira 0007, Baldoino Fonseca dos Santos Neto |
IEEE Trans. Software Eng. | 8 |
| 2017 | Smells are sensitive to developers!: on the efficiency of (un)guided customized detectionabstractCode smells indicate poor implementation choices that may hinder program comprehension and maintenance. Their informal definition allows developers to follow different heuristics to detect smells in their projects. Machine learning has been used to customize smell detection according to the developer's perception. However, such customization is not guided (i.e. constrained) to consider alternative heuristics used by developers when detecting smells. As a result, their customization might not be efficient, requiring a considerable effort to reach high effectiveness. In fact, there is no empirical knowledge yet about the efficiency of such unguided approaches for supporting developer-sensitive smell detection. This paper presents Histrategy, a guided customization technique to improve the efficiency on smell detection. Histrategy considers a limited set of detection strategies, produced from different detection heuristics, as input of a customization process. The output of the customization process consists of a detection strategy tailored to each developer. The technique was evaluated in an experimental study with 48 developers and four types of code smells. The results showed that Histrategy is able to outperform six widely adopted machine learning algorithms - used in unguided approaches - both in effectiveness and efficiency. It was also confirmed that most developers benefit from using alternative heuristics to: (i) build their tailored detection strategies, and (ii) achieve efficient smell detection. Mario Hozano, Alessandro F. Garcia 0001, Nuno Antunes, Baldoino Fonseca dos Santos Neto, Evandro de Barros Costa |
ICPC | 4 |
| 2017 | Understanding the impact of refactoring on smells: a longitudinal study of 23 software projectsabstractCode smells in a program represent indications of structural quality problems, which can be addressed by software refactoring. However, refactoring intends to achieve different goals in practice, and its application may not reduce smelly structures. Developers may neglect or end up creating new code smells through refactoring. Unfortunately, little has been reported about the beneficial and harmful effects of refactoring on code smells. This paper reports a longitudinal study intended to address this gap. We analyze how often commonly-used refactoring types affect the density of 13 types of code smells along the version histories of 23 projects. Our findings are based on the analysis of 16,566 refactorings distributed in 10 different types. Even though 79.4% of the refactorings touched smelly elements, 57% did not reduce their occurrences. Surprisingly, only 9.7% of refactorings removed smells, while 33.3% induced the introduction of new ones. More than 95% of such refactoring-induced smells were not removed in successive commits, which suggest refactorings tend to more frequently introduce long-living smells instead of eliminating existing ones. We also characterized and quantified typical refactoring-smell patterns, and observed that harmful patterns are frequent, including: (i) approximately 30% of the Move Method and Pull Up Method refactorings induced the emergence of God Class, and (ii) the Extract Superclass refactoring creates the smell Speculative Generality in 68% of the cases. Diego Cedrim, Alessandro F. Garcia 0001, Melina Mongiovi, Rohit Gheyi, Leonardo da Silva Sousa, Rafael Maiani de Mello, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Alexander Chavez |
ESEC/SIGSOFT FSE | 7 |
| 2016 | Assessing fine-grained feature dependencies
Iran Rodrigues, Márcio Ribeiro 0001, Flávio Medeiros, Paulo Borba, Baldoino Fonseca dos Santos Neto, Rohit Gheyi |
Inf. Softw. Technol. | 5 |
| 2015 | Experience report: Evaluating the effectiveness of decision trees for detecting code smellsabstractDevelopers continuously maintain software systems to adapt to new requirements and to fix bugs. Due to the complexity of maintenance tasks and the time-to-market, developers make poor implementation choices, also known as code smells. Studies indicate that code smells hinder comprehensibility, and possibly increase change- and fault-proneness. Therefore, they must be identified to enable the application of corrections. The challenge is that the inaccurate definitions of code smells make developers disagree whether a piece of code is a smell or not, consequently, making difficult creation of a universal detection solution able to recognize smells in different software projects. Several works have been proposed to identify code smells but they still report inaccurate results and use techniques that do not present to developers a comprehensive explanation how these results have been obtained. In this experimental report we study the effectiveness of the Decision Tree algorithm to recognize code smells. For this, it was applied in a dataset containing 4 open source projects and the results were compared with the manual oracle, with existing detection approaches and with other machine learning algorithms. The results showed that the approach was able to effectively learn rules for the detection of the code smells studied. The results were even better when genetic algorithms are used to pre-select the metrics to use. Lucas Amorim, Evandro de Barros Costa, Nuno Antunes, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
ISSRE | 4 |
| 2015 | AutoRefactoring: A platform to build refactoring agents
Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Viviane Torres da Silva, Christiano Braga, Carlos José Pereira de Lucena, Evandro de Barros Costa |
Expert Syst. Appl. | 1 |
| 2011 | NBDI: An Architecture for Goal-oriented Normative Agents
Baldoino Fonseca dos Santos Neto, Viviane Torres da Silva, Carlos José Pereira de Lucena |
ICAART (1) | 1 |
| 2009 | JAAF-S: A Framework to Implement Autonomic Agents Able to Deal with Web Services
Baldoino Fonseca dos Santos Neto, Andrew Diniz da Costa, Carlos José Pereira de Lucena, Viviane Torres da Silva, Manoel T. de A. Netto |
ICSOFT (1) | 1 |
| 2009 | JAAF: A Framework to Implement Self-adaptive Agents
Baldoino Fonseca dos Santos Neto, Andrew Diniz da Costa, Manoel T. de A. Netto, Viviane Torres da Silva, Carlos José Pereira de Lucena |
SEKE | 1 |