EDBT 2026 Demo / reviewers in the wild / expert
Márcio Ribeiro 0001
dblp:r/MarcioRibeiro · also Márcio de Medeiros Ribeiro
· DBLP profile ↗
68ranked-venue papers
2as first author
26since 2021 · last 2027
0000-0002-4293-4261ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 62 · 2 first-author · 24 since 2021Artificial intelligence and machine learning · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Foundation models as oracles for refactoring correctness detectionabstractAbstract Refactoring tools in popular Integrated Development Environments (IDEs) can introduce unintended behavioral changes or compilation errors, a persistent challenge that undermines developer trust in automated transformations. Traditional detection approaches rely on handcrafted preconditions, and static and dynamic analyses, yet remain limited in adaptability and can miss subtle correctness issues. This study examines the potential of foundation models to serve as oracles for detecting refactoring bugs in Java programs. We evaluate zero-shot prompting, without task-specific training, across 226 real refactoring bugs collected over more than a decade from widely used Java IDEs ( IntelliJ-IDEA , Eclipse , and NetBeans ), spanning 47 refactoring types. Our results indicate that foundation models can be effective for this task, although performance varies across models. In the first-run setting, GPT-OSS-20B achieved 80.5% accuracy, while GPT-5.4 reached 93.8%. We also evaluated other open-weight and proprietary models: Gemma-4-31B achieved the strongest result among open-weight models, and Gemini-3.1-Pro-Preview achieved the best overall result among all evaluated models. The two primary models were evaluated over five attempts, whereas the additional models were evaluated in a single-run comparison. Metamorphic testing indicates that model predictions remain largely consistent under the tested semantics-preserving perturbations, but these results should be interpreted as robustness evidence rather than as evidence against memorization or data contamination. Beyond detection accuracy, foundation models can provide short explanations that may help support developer inspection, operate across refactoring types without explicitly encoded refactoring-specific rules, and may serve as lightweight triage aids in development workflows. Our findings suggest that foundation models can complement traditional refactoring checks by flagging suspicious transformations for developer inspection. Rohit Gheyi, Rian Melo, Jonhnanthan Oliveira, Márcio Ribeiro 0001, Baldoino Fonseca dos Santos Neto |
Empir. Softw. Eng. | 4 |
| 2026 | From Legacy Designs to Vulnerability Fixes: Understanding SAST Adoption in Non-Technological Companies
Luis Henrique Vieira Amaral, Michael Schlichtig, Wagner Emanuel, Joilton Almeida, Carine Ferreira, Jerome Kempf, Rodrigo Bonifácio, Eric Bodden, Laerte Peotta, Gustavo Pinto 0001, Márcio Ribeiro 0001 |
SANER | 11 |
| 2026 | Refactoring for novices in Java: An eye tracking study on the extract vs. inline methods
José Aldo Silva da Costa, Rohit Gheyi, Silva da Costa José Júnior, Márcio Ribeiro 0001, Rodrigo Bonifácio, Hyggo Oliveira de Almeida, Ana Carla Bibiano, Alessandro F. Garcia 0001 |
J. Syst. Softw. | 4 |
| 2025 | Scaling Up: Revisiting Mining Android Sandboxes at Scale for Malware Classification (Replication Paper)abstractThe widespread use of smartphones in daily life has raised concerns about privacy and security among researchers and practitioners. Privacy issues are generally highly prevalent in mobile applications, particularly targeting the Android platform - the most popular mobile operating system. For this reason, several techniques have been proposed to identify malicious behavior in Android applications, including the Mining Android Sandbox approach (MAS approach), which aims to identify malicious behavior in repackaged Android applications (apps). However, previous empirical studies evaluated the MAS approach using a small dataset consisting of only 102 pairs of original and repackaged apps. This limitation raises questions about the external validity of their findings and whether the MAS approach can be generalized to larger datasets. To address these concerns, this paper presents the results of a replication study focused on evaluating the performance of the MAS approach regarding its capabilities of correctly classifying malware from different families. Unlike previous studies, our research employs a dataset that is an order of magnitude larger, comprising 4,076 pairs of apps covering a more diverse range of Android malware families. Surprisingly, our findings indicate a poor performance of the MAS approach for identifying malware, with the F1-score decreasing from 0.90 for the small dataset used in the previous studies to 0.54 in our more extensive dataset. Upon closer examination, we discovered that certain malware families partially account for the low accuracy of the MAS approach, which fails to classify a repackaged version of an app as malware correctly. Our findings highlight the limitations of the MAS approach, particularly when scaled, and underscore the importance of complementing it with other techniques to detect a broader range of malware effectively. This opens avenues for further discussion on addressing the blind spots that affect the accuracy of the MAS approach. Francisco Handrick da Costa, Ismael Medeiros, Leandro Oliveira 0001, João Calássio, Rodrigo Bonifácio, Krishna Narasimhan, Mira Mezini, Márcio Ribeiro 0001 |
ECOOP | 8 |
| 2025 | On the Harmfulness of Test Smells in Manual System Testing: A Controlled ExperimentabstractBackground. Test smells can pose difficulties during testing activities, such as poor maintainability, non-deterministic behavior, and incomplete verification. Existing research has extensively addressed test smells in automated software tests, but little attention has been paid to smells in natural language tests. While some research has attempted to catalog such test smells, there is a lack of investigation into their impact on the effectiveness of test cases. Aims. In this paper, we conduct a controlled experiment with 30 participants from academia and industry to examine the impact of test smells in manual test descriptions. Method. Specifically, we analyze whether the presence of two test smells, Ambiguous Test and Eager Action, result in (1) increased test execution time, (2) a higher number of steps needed to complete the tests, and (3) high divergency on the perceived success of the tests outcomes. Results. Our findings reveal that an Ambiguous Test can increase execution time by up to five times and screen flow by up to seven times. In addition, if the Eager Actions are dependent on one another, there is no increase in execution time and screen flow. Conclusions. It highlights the need for better design of manual test descriptions to improve clarity, consistency, and performance execution. Gabriela Soares, Vanessa Santos 0004, Márcio Ribeiro 0001, Luana Almeida Martins, Valeria Pontillo, Manoel Aranda III, Rohit Gheyi, Ivan do Carmo Machado, Fabio Palomba |
ESEM | 3 |
| 2025 | Evaluating the Impact of Compression Techniques on the Robustness of CNNs under Natural CorruptionsabstractCompressed deep learning models are crucial for deploying computer vision systems on resource-constrained devices. However, model compression may affect robustness, especially under natural corruption. Therefore, it is important to consider robustness evaluation while validating computer vision systems. This paper presents a comprehensive evaluation of compression techniques—quantization, pruning, and weight clustering—applied individually and in combination to convolutional neural networks (ResNet-50, VGG-19, and MobileNetV2). Using the CIFAR-10-C and CIFAR-100-C datasets, we analyze the trade-offs between robustness, accuracy, and compression ratio. Our results show that certain compression strategies not only preserve but can also improve robustness, particularly on networks with more complex architectures. Utilizing multi-objective assessment, we determine the best configurations, showing that customized technique combinations produce beneficial multi-objective results. This study provides insights into selecting compression methods for robust and efficient deployment of models in corrupted real-world environments. Itallo Patrick Castro Alves Da Silva, Emanuel Adler Medeiros Pereira, Erick A. Barboza, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
ICMLA | 5 |
| 2025 | Code Generation with Small Language Models: A Codeforces-Based StudyabstractLarge Language Models (LLMs) demonstrate capabilities in code generation, potentially boosting developer productivity. However, their adoption remains limited by high computational costs, among other factors. Small Language Models (SLMs) present a lightweight alternative. While LLMs have been evaluated on competitive programming tasks, prior work often emphasizes metrics like Elo or pass rates, neglecting failure analysis. The potential of SLMs in this space remains underexplored. In this study, we benchmark three open SLMs—Llama-3.2-3B, Gemma-3-12B, and Phi-4-14B—across 280 Codeforces problems spanning Elo ratings from 800 to 2100 and covering 36 distinct topics. All models were tasked with generating Python solutions. Phi-4-14B achieved the best SLM performance with a pass@3 of 63.6%, nearing o3-mini-high (86.8%). Combining Python and C++ outputs increased Phi-4-14B’s pass@6 to 73.6%. A qualitative analysis revealed some failures stemmed from minor implementation issues rather than reasoning flaws. Débora Souza, Rohit Gheyi, Lucas Albuquerque, Gustavo Soares, Márcio Ribeiro 0001 |
ICMLA | 5 |
| 2025 | A systematic review of fault tolerance techniques for smart city applications
Kathiani Elisa de Souza, Fabiano Cutigi Ferrari, Valter Vieira de Camargo, Márcio Ribeiro 0001, A. Jefferson Offutt |
J. Syst. Softw. | 4 |
| 2024 | A Catalog of Transformations to Remove Smells From Natural Language TestsabstractTest smells can pose difficulties during testing activities, such as poor maintainability, non-deterministic behavior, and incomplete verification. Existing research has extensively addressed test smells in automated software tests but little attention has been given to smells in natural language tests. While some research has identified and catalogued such smells, there is a lack of systematic approaches for their removal. Consequently, there is also a lack of tools to automatically identify and remove natural language test smells. This paper introduces a catalog of transformations designed to remove seven natural language test smells and a companion tool implemented using Natural Language Processing (NLP) techniques. Our work aims to enhance the quality and reliability of natural language tests during software development. The research employs a two-fold empirical strategy to evaluate its contributions. First, a survey involving 15 software testing professionals assesses the acceptance and usefulness of the catalog’s transformations. Second, an empirical study evaluates our tool to remove natural language test smells by analyzing a sample of real-practice tests from the Ubuntu OS. The results indicate that software testing professionals find the transformations valuable. Additionally, the automated tool demonstrates a good level of precision, as evidenced by a F-Measure rate of 83.70%. Manoel Aranda III, Naelson Oliveira, Elvys Soares, Márcio Ribeiro 0001, Davi Romão, Ullyanne Patriota, Rohit Gheyi, Emerson Souza, Ivan do Carmo Machado |
EASE | 4 |
| 2024 | Enhancing Recommendations of Composite Refactorings based on the PracticeabstractRefactoring is a non-trivial maintenance activity. Developers spend time and effort refactoring code to remove structural problems, i.e., code smells. Recent studies indicated that developers often apply composite refactoring (composite, for short), i.e., two or more interrelated refactorings. However, prior studies revealed that only 10% of composite refactorings are considered complete, i.e., those fully removing code smells. Many incomplete refactorings can even replace or introduce smells, requiring additional effort for their removal later in the project. Moreover, existing refactoring recommendations are not well-detailed and do not alert developers about these possible side effects. To address these gaps, we conducted a large-scale study involving more than 250k refactorings from 42 software projects, including both open-source and closed-source projects. Our goal is to investigate how the most common complete composites are combined and their side effects in the practice. Our results reveal that the current recommendation to apply Extract Method(s) with fine-grained refactoring types needs refinements. We found that certain fine-grained refactorings like Change Variable Types and Change Return Types can introduce up to 45% of Brain Methods when combined with Extract Method(s). Moreover, Ex-tract Method(s) and Move Method(s), a common recommendation to remove Feature Envy, may inadvertently introduce about 30% of Lazy Classes and approximately 70% of Data Classes. Despite these potential side effects, existing refactoring catalogs and tools' recommenders do not alert developers about these side effects. Finally, we consolidate our findings into a catalog to provide clear guidance for developers and researchers on effectively applying composite refactorings to fully remove code smells. Ana Carla Bibiano, Daniel Coutinho, Anderson G. Uchôa, Wesley K. G. Assunção, Alessandro F. Garcia 0001, Rafael Maiani de Mello, Thelma Elita Colanzi, Daniel Oliveira 0005, Audrey Vasconcelos, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
SCAM | 11 |
| 2023 | Manual Tests Do Smell! Cataloging and Identifying Natural Language Test SmellsabstractBackground: Test smells indicate potential problems in the design and implementation of automated software tests that may negatively impact test code maintainability, coverage, and reliability. When poorly described, manual tests written in natural language may suffer from related problems, which enable their analysis from the point of view of test smells. Despite the possible prejudice to manually tested software products, little is known about test smells in manual tests, which results in many open questions regarding their types, frequency, and harm to tests written in natural language. Aims: Therefore, this study aims to contribute to a catalog of test smells for manual tests. Method: We perform a two-fold empirical strategy. First, an exploratory study in manual tests of three systems: the Ubuntu Operational System, the Brazilian Electronic Voting Machine, and the User Interface of a large smartphone manufacturer. We use our findings to propose a catalog of eight test smells and identification rules based on syntactical and morphological text analysis, validating our catalog with 24 in-company test engineers. Second, using our proposals, we create a tool based on Natural Language Processing (NLP) to analyze the subject systems' tests, validating the results. Results: We observed the occurrence of eight test smells. A survey of 24 in-company test professionals showed that 80.7% agreed with our catalog definitions and examples. Our NLP-based tool achieved a precision of 92%, recall of 95%, and f-measure of 93.5%, and its execution evidenced 13,169 occurrences of our cataloged test smells in the analyzed systems. Conclusion: We contribute with a catalog of natural language test smells and novel detection strategies that better explore the capabilities of current NLP mechanisms with promising results and reduced effort to analyze tests written in different idioms. Elvys Soares, Manoel Aranda III, Naelson Oliveira, Márcio Ribeiro 0001, Rohit Gheyi, Emerson Souza, Ivan do Carmo Machado, André L. M. Santos, Baldoino Fonseca dos Santos Neto, Rodrigo Bonifácio |
ESEM | 4 |
| 2023 | The untold story of code refactoring customizations in practiceabstractRefactoring is a common software maintenance practice. The literature defines standard code modifications for each refactoring type and popular IDEs provide refactoring tools aiming to support these standard modifications. However, previous studies indicated that developers either frequently avoid using these tools or end up modifying and even reversing the code automatically refactored by IDEs. Thus, developers are forced to manually apply refactorings, which is cumbersome and error-prone. This means that refactoring support may not be entirely aligned with practical needs. The improvement of tooling support for refactoring in practice requires understanding in what ways developers tailor refactoring modifications. To address this issue, we conduct an analysis of 1,162 refactorings composed of more than 100k program modifications from 13 software projects. The results reveal that developers recurrently apply patterns of additional modifications along with the standard ones, from here on called patterns of customized refactorings. For instance, we found customized refactorings in 80.77% of the Move Method instances observed in the software projects. We also investigated the features of refactoring tools in popular IDEs and observed that most of the customization patterns are not fully supported by them. Additionally, to understand the relevance of these customizations, we conducted a survey with 40 developers about the most frequent customization patterns we found. Developers confirm the relevance of customization patterns and agree that improvements in IDE's refactoring support are needed. These observations highlight that refactoring guidelines must be updated to reflect typical refactoring customizations. Also, IDE builders can use our results as a basis to enable a more flexible application of automated refactorings. For example, developers should be able to choose which method must handle exceptions when extracting an exception code into a new method. Daniel Oliveira 0005, Wesley K. G. Assunção, Alessandro F. Garcia 0001, Ana Carla Bibiano, Márcio Ribeiro 0001, Rohit Gheyi, Baldoino Fonseca dos Santos Neto |
ICSE | 5 |
| 2023 | Automating Test-Specific Refactoring Mining: A Mixed-Method InvestigationabstractRefactoring is a practice commonly used by developers to restructure the source code without changing its external behavior. Over the last decades, the software engineering research community has been making use of mining software repository techniques to investigate refactoring under multiple perspectives, identifying properties and impact of this practice on source code quality, other than using refactoring data coming from software repositories to build automated recommendation systems. While the current state of the art proposes various automated tools to mine refactoring data, there is still a lack of instruments that may help researchers when mining test-specific refactoring data. The availability of those instruments may enable additional, specialized techniques to support developers while refactoring test code. In this paper, we introduce an approach that extends REFACTORINGMINER-a well-established refactoring mining tool having high precision and recall scores- and is able to detect seven test-specific refactoring operations. We perform mixed-method research to assess capabilities and usefulness of the approach. First, we compare the test-specific refactoring data extracted by the approach against an oracle of 375 test-specific refactorings. Second, we engage with 15 software engineering researchers and apply a technology acceptance model to investigate how they would benefit from our approach. The key results of the study show that our approach reaches 100% and 92.5% of precision and recall scores, respectively. In addition, the approach is considered useful and suitable for various research tasks, including the definition of novel learning models able to recommend test-specific refactoring actions. Luana Almeida Martins, Heitor A. X. Costa, Márcio Ribeiro 0001, Fabio Palomba, Ivan do Carmo Machado |
SCAM | 3 |
| 2023 | Seeing confusion through a new lens: on the impact of atoms of confusion on novices' code comprehension
José Aldo Silva da Costa, Rohit Gheyi, Fernando Castor Filho, Pablo Roberto Fernandes de Oliveira, Márcio Ribeiro 0001, Baldoino Fonseca dos Santos Neto |
Empir. Softw. Eng. | 5 |
| 2023 | Towards a better understanding of the mechanics of refactoring detection tools
Jonhnanthan Oliveira, Rohit Gheyi, Leopoldo Teixeira, Márcio Ribeiro 0001, Osmar Leandro, Baldoino Fonseca dos Santos Neto |
Inf. Softw. Technol. | 4 |
| 2023 | An Investigation of confusing code patterns in JavaScript
Adriano Torres, Caio Oliveira, Márcio Vinicius Okimoto, Diego Marcilio, Pedro Queiroga, Fernando Castor Filho, Rodrigo Bonifácio, Edna Dias Canedo, Márcio Ribeiro 0001, Eduardo Monteiro |
J. Syst. Softw. | 9 |
| 2023 | Refactoring Test Smells With JUnit 5: Why Should Developers Keep Up-to-Date?abstractTest smells are symptoms in the test code that indicate possible design or implementation problems. Previous research demonstrated their harmfulness and the developers’ acknowledgment of test smells’ effects, prevention, and refactoring strategies. Test automation frameworks are constantly evolving, and the JUnit, one of the most used ones for Java projects, has its version 5 available since late 2017. However, we do not know the extent to which developers use the newly introduced features and whether such features indeed help refactor existing test code to remove test smells. This article conducts a mixed-method study investigation to minimize these knowledge gaps. Our study consists of three parts. First, we evaluate the usage of this framework and its features by analyzing the source code of 485 popular Java open-source projects on GitHub that use JUnit. We found that 15.9% of these projects use the JUnit 5 library. We also found that, from 17 new features detected in use, only 3 (i.e., 17.6%) are responsible for more than 70% of usages, limiting optimized propositions to test code creation and maintenance. Second, after identifying features in the JUnit 5 framework that could be considered to test smells removal and prevention, we use these features to propose novel refactorings. In particular, we present refactorings based on 7 introduced JUnit 5 features that help to remove 13 test smells, such as Assertion Roulette, Test Code Duplication, and Conditional Test Logic. Third, to evaluate our refactorings with the opinions of experienced developers, we (i) survey 212 developers for their preferences and comments about our refactorings, corroborating the benefits of our proposals and raising community feedback on JUnit 5 features, and (ii) we refactor actual test code from popular GitHub Java projects and submit 38 Pull Requests, reaching a 94% acceptance rate among respondents. As implications of our study, we alert the software testing community (i.e., practitioners and researchers) to the need to study the JUnit 5 features to effectively remove and prevent test smells. To better assist this process, we give directions on how test smells can be refactored using such features. Elvys Soares, Márcio Ribeiro 0001, Rohit Gheyi, Guilherme Amaral, André L. M. Santos |
IEEE Trans. Software Eng. | 2 |
| 2022 | Lint-Based Warnings in Python Code: Frequency, Awareness and RefactoringabstractPython is a popular programming language characterized by its simple syntax and easy learning curve. Like many languages, Python has a set of best practices that should be followed to avoid bugs and improve other quality attributes (such as maintenance and readability). In this context, non-compliance to these practices can be detected by using linting tools. Previous work conducted studies to better understand the frequency of a class of problems that can be found using Python linters: warnings, here named as lint-based warnings. However, they either rely on small datasets or focus on few domains, such as machine learning or web-systems projects. In this paper, we provide a mixed-method study where we analyze the frequency of six lint-based warnings in 1,119 different open-source general-purpose Python projects. To go further, we also conduct a survey to check whether developers are aware of the lint-based warnings we study here. In particular, we intend to check whether they are able to identify the six lint-based warnings. To remove the lint-based warnings, we suggest the application of simple refactorings. Last but not least, we evaluate the suggestions by submitting pull requests to remove lint-based warnings from open-source projects. Our results show that 39% of the 1,119 projects have at least one lint-based warning. After analyzing the survey data, we also show that developers prefer Python code without lint-based warnings. Regarding the pull requests, we achieve a 71.8% of acceptance rate. Naelson Oliveira, Márcio Ribeiro 0001, Rodrigo Bonifácio, Rohit Gheyi, Igor Scaliante Wiese, Baldoino Fonseca dos Santos Neto |
SCAM | 2 |
| 2022 | Developers' perception matters: machine learning to detect developer-sensitive smells
Daniel Oliveira 0005, Wesley K. G. Assunção, Alessandro F. Garcia 0001, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
Empir. Softw. Eng. | 5 |
| 2022 | Developers' viewpoints to avoid bug-introducing changes
Jairo Souza, Rodrigo Lima 0002, Baldoino Fonseca dos Santos Neto, Bruno Cartaxo, Márcio Ribeiro 0001, Gustavo Pinto 0001, Rohit Gheyi, Alessandro F. Garcia 0001 |
Inf. Softw. Technol. | 5 |
| 2022 | Exploring the use of static and dynamic analysis to improve the performance of the mining sandbox approach for android malware identification
Francisco Handrick da Costa, Ismael Medeiros, Thales Menezes, João Victor da Silva, Ingrid Lorraine da Silva, Rodrigo Bonifácio, Krishna Narasimhan, Márcio Ribeiro 0001 |
J. Syst. Softw. | 8 |
| 2021 | What Evidence We Would Miss If We Do Not Use Grey Literature?abstractContext: Multivocal Literature Reviews (MLR) search for evidence in both Traditional Literature (TL) and Grey Literature (GL). Despite the growing interest in MLR-based studies, the literature assessing how GL has contributed to MLR studies is still scarce. Objective: This research aims to assess how the use of GL contributed to MLR studies. By contributing, we mean, understanding to what extent GL is providing evidence that is indeed used by an MLR to answer its research question. Method: We start by conducting a tertiary study to identify MLR studies published between 2017 and 2019, selecting nine of them. We then identified the GL used in these studies and assessed to what extent the GLs are providing evidence that help these studies to answer their research questions. Results: Our analysis identified that 1) GL provided evidence not found in TL, 2) most of the GL sources were used to provide recommendations to solve problems, explain a topic, and classify the findings, and 3) 19 different GL types were used in the studies; these GLs were mainly produced by SE practitioners (including blog posts, slides presentations, or project descriptions). Conclusions: We evidence how GL contributed to MLR studies. We observed that if these GLs were not included in the MLR, several findings would have been omitted or weakened. We also described the challenges involved when conducting this investigation, along with potential ways to deal with them, which may help future SE researchers. Fernando Kamei, Gustavo Pinto 0001, Igor Scaliante Wiese, Márcio Ribeiro 0001, Sérgio Soares |
ESEM | 4 |
| 2021 | Look Ahead! Revealing Complete Composite Refactorings and their Smelliness EffectsabstractRecent studies have revealed that developers often apply composite refactorings (or, simply, composites). A composite consists of two or more interrelated refactorings applied together. Previous studies investigated the effect of composites on code smells. A composite is considered “complete” whenever it completely removes one target code smell. They proposed descriptions of complete composites with recommendations to remove certain code smell types, such as Long Methods and Feature Envies. These studies also present different recommendations to remove the same code smell type. However, these studies: (i) are limited to composites only consisting of a small subset of Fowler's refactoring types, (ii) do not detail the scenarios in which each recommendation can be applied to remove the code smell, and (iii) fail in reporting possible side effects of the described composites, such as adversely introducing certain smell types. This paper aims to cover these limitations by performing a systematic analysis of 618 complete composites on removing four common smell types identified in 20 software projects. Our results indicated that: (i) 64% complete composites consisted of refactoring types not covered by existing descriptions of complete composites, and (ii) 36% complete composites formed by Extract Methods can introduce Feature Envies and Intensive Couplings. This information is not documented by existing descriptions, and it can alert developers about alternatives to remove Feature Envy, mainly in methods that are fully envious. These results suggest existing descriptions of complete composites should be either revisited or enhanced to explicitly highlight known side effects. We present a catalog of composites with details about side effects, recommendations to remove or minimize them, and some scenarios in which each recommendation can be applied to remove the code smell. Our catalog can be useful to improve existing tooling support for refactorings, such as IDEs, informing about possible side effects when refactorings are composed. Ana Carla Bibiano, Wesley K. G. Assunção, Daniel Coutinho, Kleber Santos, Vinícius Soares, Rohit Gheyi, Alessandro F. Garcia 0001, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Daniel Oliveira 0005, Caio Barbosa, João Lucas Marques, Anderson Oliveira |
ICSME | 9 |
| 2021 | Evaluating refactorings for disciplining #ifdef annotations: An eye tracking study with novices
José Aldo Silva da Costa, Rohit Gheyi, Márcio Ribeiro 0001, Sven Apel, Vander Alves, Baldoino Fonseca dos Santos Neto, Flávio Medeiros, Alessandro F. Garcia 0001 |
Empir. Softw. Eng. | 3 |
| 2021 | Identifying method-level mutation subsumption relations using Z3
Rohit Gheyi, Márcio Ribeiro 0001, Beatriz Souza, Marcio Augusto Guimarães, Leonardo Fernandes, Marcelo d'Amorim, Vander Alves, Leopoldo Teixeira, Baldoino Fonseca dos Santos Neto |
Inf. Softw. Technol. | 2 |
| 2021 | Grey Literature in Software Engineering: A critical review
Fernando Kamei, Igor Scaliante Wiese, Crescencio Rodrigues Lima Neto, Ivanilton Polato, Vilmar Nepomuceno, Waldemar Ferreira, Márcio Ribeiro 0001, Carolline Pena, Bruno Cartaxo, Gustavo Pinto 0001, Sérgio Soares |
Inf. Softw. Technol. | 7 |
| 2020 | Is Exceptional Behavior Testing an Exception?: An Empirical Assessment Using Java Automated TestsabstractSoftware testing is a crucial activity to check the internal quality of a software. During testing, developers often create tests for the normal behavior of a particular functionality (e.g., was this file properly uploaded to the cloud?). However, little is known whether developers also create tests for the exceptional behavior (e.g., what happens if the network fails during the file upload?). To minimize this knowledge gap, in this paper we design and perform a mixed-method study to understand how 417 open source Java projects are testing the exceptional behavior using the JUnit and TestNG frameworks, and the AssertJ library. We found that 254 (60.91%) projects have at least one test method dedicated to test the exceptional behavior. We also found that the number of test methods for exceptional behavior with respect to the total number of test methods lies between 0% and 10% in 317 (76.02%) projects. Also, 239 (57.31%) projects test only up to 10% of the used exceptions in the System Under Test (SUT). When it comes to mobile apps, we found that, in general, developers pay less attention to exceptional behavior tests when compared to desktop/server and multi-platform developers. In general, we found more test methods covering custom exceptions (the ones created in the own project) when compared to standard exceptions available in the Java Development Kit (JDK) or in third-party libraries. To triangulate the results, we conduct a survey with 66 developers from the projects we study. In general, the survey results confirm our findings. In particular, the majority of the respondents agrees that developers often neglect exceptional behavior tests. As implications, our numbers might be important to alert developers that more effort should be placed on creating tests for the exceptional behavior. Francisco Dalton, Márcio Ribeiro 0001, Gustavo Pinto 0001, Leonardo Fernandes, Rohit Gheyi, Baldoino Fonseca dos Santos Neto |
EASE | 2 |
| 2020 | On the Performance and Adoption of Search-Based Microservice Identification with toMicroservicesabstractThe expensive maintenance of legacy systems leads companies to migrate such systems to microservice architectures. This migration requires the identification of system's legacy parts to become microservices. However, the successful identification of microservices, which are promising to be adoptable in practice, requires the simultaneous satisfaction of many criteria, such as coupling, cohesion, reuse and communication overhead. Search-based microservice identification has been recently investigated to address this problem. However, state-of-the-art search-based approaches are limited as they only consider one or two criteria (namely cohesion and coupling), possibly not fulfilling the practical needs of developers. To overcome these limitations, we propose toMicroservices, a many-objective search-based approach that considers five criteria, the most cited by practitioners in recent studies. Our approach was evaluated in a real-life industrial legacy system undergoing a microservice migration process. The performance of toMicroservices was quantitatively compared to a baseline. We also gathered qualitative evidence based on developers' perceptions, who judged the adoptability of the recommended microservices. The results show that our approach is both: (i) very similar to the most recent proposed approach on optimizing the traditional criteria of coupling and cohesion, but (ii) much better when taking into account all the five criteria. Finally, most of the microservice candidates were considered adoptable by practitioners. Alessandro F. Garcia 0001, Thelma Elita Colanzi, Wesley K. G. Assunção, Juliana Alves Pereira, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Maria Julia de Lima, Carlos José Pereira de Lucena |
ICSME | 7 |
| 2020 | Optimizing Mutation Testing by Discovering Dynamic Mutant Subsumption RelationsabstractOne recent promising direction on reducing costs of mutation analysis is to identify redundant mutations, i.e., mutations that are subsumed by some other mutations. Previous works found out redundant mutants manually through the truth table. Although the idea is promising, it can only be applied for logical and relational operators. In this paper, we propose an approach to discover redundancy in mutations through dynamic subsumption relations among mutants. We focus on subsumption relations among mutations of an expression or statement, named here as “mutation target:” By focusing on targets and relying on automatic test generation tools, we define subsumption relations for dozens of mutation targets in which the MUJAVA tool can apply mutations. We then implemented these relations in a tool, named MUJAVA-M, that generates a reduced set of mutants for each target, avoiding redundant mutants. We evaluated MUJAVA and MUJAVA-M using classes of five open-source projects. As results, we analyze 2,341 occurrences of 32 mutation targets in 168 classes. MUJAVA-M generates less mutants (on average 64.43% less) with 100% of effectiveness in 20 out of 32 targets and more than 95% in 29 out of 32 mutation targets. MUJAVA- M also reduced the time to execute the test suites against the mutants in 52.53% on average, considering the full mutation analysis process. Marcio Augusto Guimarães, Leonardo Fernandes, Márcio Ribeiro 0001, Marcelo d'Amorim, Rohit Gheyi |
ICST | 3 |
| 2020 | How Does Incomplete Composite Refactoring Affect Internal Quality Attributes?abstractProgram refactoring consists of code changes applied to improve the internal structure of a program and, as a consequence, its comprehensibility. Recent studies indicate that developers often perform composite refactorings, i.e., a set of two or more interrelated single refactorings. Recent studies also recommend certain patterns of composite refactorings to fully remove poor code structures, i.e, code smells, thus further improving the program comprehension. However, other recent studies report that composite refactorings often fail to fully remove code smells. Given their failure to achieve this purpose, these composite refactorings are considered incomplete, i.e, they are not able to entirely remove a smelly structure. Unfortunately, there is no study providing an in-depth analysis of the incompleteness nature of many composites and their possibly partial impact on improving, maybe decreasing, internal quality attributes. This paper identifies the most common forms of incomplete composites, and their effect on quality attributes, such as coupling and cohesion, which are known to have an impact on program comprehension. We analyzed 353 incomplete composite refactorings in 5 software projects, two common code smells (Feature Envy and God Class), and four internal quality attributes. Our results reveal that incomplete composite refactorings with at least one Extract Method are often (71%) applied without Move Methods on smelly classes. We have also found that most incomplete composite refactorings (58%) tended to at least maintain the internal structural quality of smelly classes, thereby not causing more harm to program comprehension. We also discuss the implications of our findings to the research and practice of composite refactoring. Ana Carla Bibiano, Vinícius Soares, Daniel Coutinho, Eduardo Fernandes, João Lucas Correia, Kleber Santos, Anderson Oliveira, Alessandro F. Garcia 0001, Rohit Gheyi, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Caio Barbosa, Daniel Oliveira 0005 |
ICPC | 11 |
| 2020 | Refactoring from 9 to 5? What and When Employees and Volunteers Contribute to OSSabstractIn this paper we characterize the contributions made by employees (developers that work for GitHub, the company) and volunteers (developers that use GitHub, the platform) to OSS projects maintained by GitHub (the company) on GitHub (the platform). By mining activities performed in five well-known company-owned OSS projects, we investigate what they do and when they do it. We found that the majority of the volunteers' contributions are related to reengineering (e.g., refactoring), while employees focus more on management (e.g., documentation). When it comes to the working hours, we found that contributions are made mostly from 9am-5pm, even for the volunteers. Luiz Felipe Dias, Caio Barbosa, Gustavo Pinto 0001, Igor Steinmacher, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Christoph Treude, Daniel Alencar da Costa |
VL/HCC | 6 |
| 2020 | On Relating Technical, Social Factors, and the Introduction of BugsabstractAs collaborative coding environments make it easier to contribute to software projects, the number of developers involved in these projects keeps increasing. This increase makes it more difficult for code reviewers to deal with buggy contributions. Collaborative environments like GitHub provide a rich source of data on developers' contributions. Such data can be used to extract information about developers regarding technical (e.g., their experience) and social (e.g., their interactions) factors. Recent studies analyzed the influence of these factors on different activities of software development. However, there is still room for improvement on the relation between these factors and the introduction of bugs. We present a broader study, including 8 projects from different domains and 6,537 bug reports, on relating five technical, three social factors, and the introduction of bugs. The results indicate that technical and social factors can discriminate between buggy and clean commits. But, the technical factors are more determining than social ones. Particularly, the developers' habits of not following technical contribution norms and the developer's commit bugginess are associated with an increase on commit bugginess. On the other hand, project's establishment, ownership level of developers' commit, and social influence are related to a lower chance of introducing bugs. Filipe Falcão, Caio Barbosa, Baldoino Fonseca dos Santos Neto, Alessandro F. Garcia 0001, Márcio Ribeiro 0001, Rohit Gheyi |
SANER | 5 |
| 2020 | Mutating code annotations: An empirical evaluation on Java and C# programs
Pedro Pinheiro, José Carlos Viana, Márcio Ribeiro 0001, Leonardo Fernandes, Fabiano Cutigi Ferrari, Rohit Gheyi, Baldoino Fonseca dos Santos Neto |
Sci. Comput. Program. | 3 |
| 2019 | Software Engineering Research Community Viewpoints on Rapid ReviewsabstractBackground: One of the most important current challenges of Software Engineering (SE) research is to provide relevant evidence to practice. In health related fields, Rapid Reviews (RRs) have shown to be an effective method to achieve that goal. However, little is known about how the SE research community perceives the potential applicability of RRs. Aims: The goal of this study is to understand the SE research community viewpoints towards the use of RRs as a means to provide evidence to practitioners. Method: To understand their viewpoints, we invited 37 researchers to analyze 50 opinion statements about RRs, and rate them according to what extent they agree with each statement. Q-Methodology was employed to identify the most salient viewpoints, represented by the so called factors. Results: Four factors were identified: Factor A groups undecided researchers that need more evidence before using RRs; Researchers grouped in Factor B are generally positive about RRs, but highlight the need to define minimum standards; Factor C researchers are more skeptical and reinforce the importance of high quality evidence; Researchers aligned to Factor D have a pragmatic point of view, considering RRs can be applied based on the context and constraints faced by practitioners. Conclusions: In conclusion, although there are opposing viewpoints, there are also some common grounds. For example, all viewpoints agree that both RRs and Systematic Reviews can be poorly or well conducted. Bruno Cartaxo, Gustavo Pinto 0001, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Pedro Pinheiro, Maria Teresa Baldassarre, Sérgio Soares |
ESEM | 4 |
| 2019 | Java reflection API: revealing the dark side of the mirrorabstractDevelopers of widely used Java Virtual Machines (JVMs) implement and test the Java Reflection API based on a Javadoc, which is specified using a natural language. However, there is limited knowledge on whether Java Reflection API developers are able to systematically reveal i) underdetermined specifications; and ii) non-conformances between their implementation and the Javadoc. Moreover, current automatic test suite generators cannot be used to detect them. To better understand the problem, we analyze test suites of two widely used JVMs, and we conduct a survey with 130 developers who use the Java Reflection API to see whether the Javadoc impacts on their understanding. We also propose a technique to detect underdetermined specifications and non-conformances between the Javadoc and the implementations of the Java Reflection API. It automatically creates test cases, and executes them using different JVMs. Then, we manually execute some steps to identify underdetermined specifications and to confirm whether a non-conformance candidate is indeed a bug. We evaluate our technique in 439 input programs. Our technique identifies underdetermined specification and non-conformance candidates in 32 Java Reflection API public methods of 7 classes. We report underdetermined specification candidates in 12 Java Reflection API methods. Java Reflection API specifiers accept 3 underdetermined specification candidates (25%). We also report 24 non-conformance candidates to Eclipse OpenJ9 JVM, and 7 to Oracle JVM. Eclipse OpenJ9 JVM developers accept and fix 21 candidates (87.5%), and Oracle JVM developers accept 5 and fix 4 non-conformance candidates. Felipe Pontes, Rohit Gheyi, Sabrina Souto, Alessandro F. Garcia 0001, Márcio Ribeiro 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2019 | An investigation of misunderstanding code patterns in C open-source software projects
Flávio Medeiros, Gabriel Lima, Guilherme Amaral, Sven Apel, Christian Kästner, Márcio Ribeiro 0001, Rohit Gheyi |
Empir. Softw. Eng. | 6 |
| 2019 | Revisiting the refactoring mechanics
Jonhnanthan Oliveira, Rohit Gheyi, Melina Mongiovi, Gustavo Soares, Márcio Ribeiro 0001, Alessandro F. Garcia 0001 |
Inf. Softw. Technol. | 5 |
| 2019 | A systematic literature review of techniques and metrics to reduce the cost of mutation testing
Alessandro Viola Pizzoleto, Fabiano Cutigi Ferrari, A. Jefferson Offutt, Leonardo Fernandes, Márcio Ribeiro 0001 |
J. Syst. Softw. | 5 |
| 2018 | Measuring effectiveness of sample-based product-line testingabstractRecent research on quality assurance (QA) of configurable software systems (e.g., software product lines) proposes different analysis strategies to cope with the inherent complexity caused by the well-known combinatorial-explosion problem. Those strategies aim at improving efficiency of QA techniques like software testing as compared to brute-force configuration-by-configuration analysis. Sampling constitutes one of the most established strategies, defining criteria for selecting a drastically reduced, yet sufficiently diverse subset of software configurations considered during QA. However, finding generally accepted measures for assessing the impact of sample-based analysis on the effectiveness of QA techniques is still an open issue. We address this problem by lifting concepts from single-software mutation testing to configurable software. Our framework incorporates a rich collection of mutation operators for product lines implemented in C to measure mutation scores of samples, including a novel family-based technique for product-line mutation detection. Our experimental results gained from applying our tool implementation to a collection of subject systems confirms the widely-accepted assumption that pairwise sampling constitutes the most reasonable efficiency/effectiveness trade-off for sample-based product-line testing. Sebastian Ruland, Lars Luthmann, Johannes Bürdek, Sascha Lity, Thomas Thüm, Malte Lochau, Márcio Ribeiro 0001 |
GPCE | 7 |
| 2018 | A change-aware per-file analysis to compile configurable systems with #ifdefs
Larissa Braz, Rohit Gheyi, Melina Mongiovi, Márcio Ribeiro 0001, Flávio Medeiros, Leopoldo Teixeira, Sabrina Souto |
Comput. Lang. Syst. Struct. | 4 |
| 2018 | Variability Bugs in Highly Configurable Systems: A Qualitative AnalysisabstractVariability-sensitive verification pursues effective analysis of the exponentially many variants of a program family. Several variability-aware techniques have been proposed, but researchers still lack examples of concrete bugs induced by variability, occurring in real large-scale systems. A collection of real world bugs is needed to evaluate tool implementations of variability-sensitive analyses by testing them on real bugs. We present a qualitative study of 98 diverse variability bugs (i.e., bugs that occur in some variants and not in others) collected from bug-fixing commits in the Linux, Apache, BusyBox, and Marlin repositories. We analyze each of the bugs, and record the results in a database. For each bug, we create a self-contained simplified version and a simplified patch, in order to help researchers who are not experts on these subject studies to understand them, so that they can use these bugs for evaluation of their tools. In addition, we provide single-function versions of the bugs, which are useful for evaluating intra-procedural analyses. A web-based user interface for the database allows to conveniently browse and visualize the collection of bugs. Our study provides insights into the nature and occurrence of variability bugs in four highly-configurable systems implemented in C/C++, and shows in what ways variability hinders comprehension and the uncovering of software bugs. Iago Abal, Jean Melo, Stefan Stanciulescu, Claus Brabrand, Márcio Ribeiro 0001, Andrzej Wasowski |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2018 | Discipline Matters: Refactoring of Preprocessor Directives in the #ifdef HellabstractThe C preprocessor is used in many C projects to support variability and portability. However, researchers and practitioners criticize the C preprocessor because of its negative effect on code understanding and maintainability and its error proneness. More importantly, the use of the preprocessor hinders the development of tool support that is standard in other languages, such as automated refactoring. Developers aggravate these problems when using the preprocessor in undisciplined ways (e.g., conditional blocks that do not align with the syntactic structure of the code). In this article, we proposed a catalogue of refactorings and we evaluated the number of application possibilities of the refactorings in practice, the opinion of developers about the usefulness of the refactorings, and whether the refactorings preserve behavior. Overall, we found 5,670 application possibilities for the refactorings in 63 real-world C projects. In addition, we performed an online survey among 246 developers, and we submitted 28 patches to convert undisciplined directives into disciplined ones. According to our results, 63 percent of developers prefer to use the refactored (i.e., disciplined) version of the code instead of the original code with undisciplined preprocessor usage. To verify that the refactorings are indeed behavior preserving, we applied them to more than 36 thousand programs generated automatically using a model of a subset of the C language, running the same test cases in the original and refactored programs. Furthermore, we applied the refactorings to three real-world projects: BusyBox, OpenSSL, and SQLite. This way, we detected and fixed a few behavioral changes, 62 percent caused by unspecified behavior in the C programming language. Flávio Medeiros, Márcio Ribeiro 0001, Rohit Gheyi, Sven Apel, Christian Kästner, Bruno Ferreira 0007, Baldoino Fonseca dos Santos Neto |
IEEE Trans. Software Eng. | 2 |
| 2018 | Detecting Overly Strong Preconditions in Refactoring EnginesabstractRefactoring engines may have overly strong preconditions preventing developers from applying useful transformations. We find that 32 percent of the Eclipse and JRRT test suites are concerned with detecting overly strong preconditions. In general, developers manually write test cases, which is costly and error prone. Our previous technique detects overly strong preconditions using differential testing. However, it needs at least two refactoring engines. In this work, we propose a technique to detect overly strong preconditions in refactoring engines without needing reference implementations. We automatically generate programs and attempt to refactor them. For each rejected transformation, we attempt to apply it again after disabling the preconditions that lead the refactoring engine to reject the transformation. If it applies a behavior preserving transformation, we consider the disabled preconditions overly strong. We evaluate 10 refactorings of Eclipse and JRRT by generating 154,040 programs. We find 15 overly strong preconditions in Eclipse and 15 in JRRT. Our technique detects 11 bugs that our previous technique cannot detect while missing 5 bugs. We evaluate the technique by replacing the programs generated by JDolly with the input programs of Eclipse and JRRT test suites. Our technique detects 14 overly strong preconditions in Eclipse and 4 in JRRT. Melina Mongiovi, Rohit Gheyi, Gustavo Soares, Márcio Ribeiro 0001, Paulo Borba, Leopoldo Teixeira |
IEEE Trans. Software Eng. | 4 |
| 2017 | Avoiding useless mutantsabstractMutation testing is a program-transformation technique that injects artificial bugs to check whether the existing test suite can detect them. However, the costs of using mutation testing are usually high, hindering its use in industry. Useless mutants (equivalent and duplicated) contribute to increase costs. Previous research has focused mainly on detecting useless mutants only after they are generated and compiled. In this paper, we introduce a strategy to help developers with deriving rules to avoid the generation of useless mutants. To use our strategy, we pass as input a set of programs. For each program, we also need a passing test suite and a set of mutants. As output, our strategy yields a set of useless mutants candidates. After manually confirming that the mutants classified by our strategy as "useless" are indeed useless, we derive rules that can avoid their generation and thus decrease costs. To the best of our knowledge, we introduce 37 new rules that can avoid useless mutants right before their generation. We then implement a subset of these rules in the MUJAVA mutation testing tool. Since our rules have been derived based on artificial and small Java programs, we take our MUJAVA version embedded with our rules and execute it in industrial-scale projects. Our rules reduced the number of mutants by almost 13% on average. Our results are promising because (i) we avoid useless mutants generation; (ii) our strategy can help with identifying more rules in case we set it to use more complex Java programs; and (iii) our MUJAVA version has only a subset of the rules we derived. Leonardo Fernandes, Márcio Ribeiro 0001, Rohit Gheyi, Melina Mongiovi, André L. M. Santos, Ana Cavalcanti 0001, Fabiano Cutigi Ferrari, José Carlos Maldonado |
GPCE | 2 |
| 2017 | The discipline of preprocessor-based annotations does #ifdef TAG n't #endif matterabstractThe C preprocessor is a simple, effective, and language-independent tool. Developers use the preprocessor in practice to deal with portability and variability issues. Despite the widespread usage, the C preprocessor suffers from severe criticism, such as negative effects on code understandability and maintainability. In particular, these problems may get worse when using undisciplined annotations, i.e., when a preprocessor directive encompasses only parts of C syntactical units. Nevertheless, despite the criticism and guidelines found in systems like Linux to avoid undisciplined annotations, the results of a previous controlled experiment indicated that the discipline of annotations has no influence on program comprehension and maintenance. To better understand whether developers care about the discipline of preprocessor-based annotations and whether they can really influence on maintenance tasks, in this paper we conduct a mixed-method research involving two studies. In the first one, we identify undisciplined annotations in 110 open-source C/C++ systems of different domains, sizes, and popularity GitHub metrics. We then refactor the identified undisciplined annotations to make them disciplined. Right away, we submit pull requests with our code changes. Our results show that almost two thirds of our pull requests have been accepted and are now merged. In the second study, we conduct a controlled experiment. We have several differences with respect to the aforementioned one, such as blocking of cofounding effects and more replicas. We have evidences that maintaining undisciplined annotations is more time consuming and error prone, representing a different result when compared to the previous experiment. Overall, we conclude that undisciplined annotations should not be neglected. Romero Malaquias, Márcio Ribeiro 0001, Rodrigo Bonifácio, Eduardo Monteiro, Flávio Medeiros, Alessandro F. Garcia 0001, Rohit Gheyi |
ICPC | 2 |
| 2017 | Understanding the impact of refactoring on smells: a longitudinal study of 23 software projectsabstractCode smells in a program represent indications of structural quality problems, which can be addressed by software refactoring. However, refactoring intends to achieve different goals in practice, and its application may not reduce smelly structures. Developers may neglect or end up creating new code smells through refactoring. Unfortunately, little has been reported about the beneficial and harmful effects of refactoring on code smells. This paper reports a longitudinal study intended to address this gap. We analyze how often commonly-used refactoring types affect the density of 13 types of code smells along the version histories of 23 projects. Our findings are based on the analysis of 16,566 refactorings distributed in 10 different types. Even though 79.4% of the refactorings touched smelly elements, 57% did not reduce their occurrences. Surprisingly, only 9.7% of refactorings removed smells, while 33.3% induced the introduction of new ones. More than 95% of such refactoring-induced smells were not removed in successive commits, which suggest refactorings tend to more frequently introduce long-living smells instead of eliminating existing ones. We also characterized and quantified typical refactoring-smell patterns, and observed that harmful patterns are frequent, including: (i) approximately 30% of the Move Method and Pull Up Method refactorings induced the emergence of God Class, and (ii) the Extract Superclass refactoring creates the smell Speculative Generality in 68% of the cases. Diego Cedrim, Alessandro F. Garcia 0001, Melina Mongiovi, Rohit Gheyi, Leonardo da Silva Sousa, Rafael Maiani de Mello, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Alexander Chavez |
ESEC/SIGSOFT FSE | 8 |
| 2017 | An idiom to represent data types in Alloy
Rohit Gheyi, Paulo Borba, Augusto Sampaio 0001, Márcio Ribeiro 0001 |
Inf. Softw. Technol. | 4 |
| 2016 | A change-centric approach to compile configurable systems with #ifdefsabstractConfigurable systems typically use #ifdefs to denote variability. Generating and compiling all configurations may be time-consuming. An alternative consists of using variability-aware parsers, such as TypeChef. However, they may not scale. In practice, compiling the complete systems may be costly. Therefore, developers can use sampling strategies to compile only a subset of the configurations. We propose a change-centric approach to compile configurable systems with #ifdefs by analyzing only configurations impacted by a code change (transformation). We implement it in a tool called CHECKCONFIGMX, which reports the new compilation errors introduced by the transformation. We perform an empirical study to evaluate 3,913 transformations applied to the 14 largest files of BusyBox, Apache HTTPD, and Expat configurable systems. CHECKCONFIGMX finds 595 compilation errors of 20 types introduced by 41 developers in 214 commits (5.46% of the analyzed transformations). In our study, it reduces by at least 50% (an average of 99%) the effort of evaluating the analyzed transformations by comparing with the exhaustive approach without considering a feature model. CHECKCONFIGMX may help developers to reduce compilation effort to evaluate fine-grained transformations applied to configurable systems with #ifdefs. Larissa Braz, Rohit Gheyi, Melina Mongiovi, Márcio Ribeiro 0001, Flávio Medeiros, Leopoldo Teixeira |
GPCE | 4 |
| 2016 | A comparison of 10 sampling algorithms for configurable systemsabstractAlmost every software system provides configuration options to tailor the system to the target platform and application scenario. Often, this configurability renders the analysis of every individual system configuration infeasible. To address this problem, researchers have proposed a diverse set of sampling algorithms. We present a comparative study of 10 state-of-the-art sampling algorithms regarding their fault-detection capability and size of sample sets. The former is important to improve software quality and the latter to reduce the time of analysis. In a nutshell, we found that sampling algorithms with larger sample sets are able to detect higher numbers of faults, but simple algorithms with small sample sets, such as most-enabled-disabled, are the most efficient in most contexts. Furthermore, we observed that the limiting assumptions made in previous work influence the number of detected faults, the size of sample sets, and the ranking of algorithms. Finally, we have identified a number of technical challenges when trying to avoid the limiting assumptions, which questions the practicality of certain sampling algorithms. Flávio Medeiros, Christian Kästner, Márcio Ribeiro 0001, Rohit Gheyi, Sven Apel |
ICSE | 3 |
| 2016 | Product-line maintenance with emergent contract interfacesabstractA software product line evolves whenever one of its products need to evolve. Maintenance of preprocessor-based product lines is a difficult task, as changes to the code base may unintentionally influence the behavior of uninvolved products. Hence, developers should be supported during maintenance. We present emergent contract interfaces to make product-line development more efficient and less error-prone. The key idea is that for a given maintenance point (i.e., an assignment), we calculate (a) features in the source code that may be affected and (b) assertions based on contracts defined in the code base. By means of a controlled experiment, we provide empirical evidence regarding efficiency and error-avoidance with emergent contract interfaces. Thomas Thüm, Márcio Ribeiro 0001, Reimar Schröter, Janet Siegmund, Francisco Dalton |
SPLC | 2 |
| 2016 | Assessing Idioms for a Flexible Feature Binding TimeabstractIn software product lines development, it is sometimes important to provide a flexible binding time for features such that developers can choose between static or dynamic feature activation. For example, software products designed for devices with constrained resources may use a static binding time to avoid the performance overhead introduced by dynamic binding time activation. However, other devices can exploit binding time flexibility to support products with a dynamic binding time for some of their features. To implement this kind of flexibility in a modular way, we can define AspectJ-based idioms. Researchers have proposed Edicts, an idiom based on AspectJ and design patterns. In this article, we argue that this idiom leads to an increase in code duplication, scattering, tangling and size, which can hamper code reuse, maintenance and understanding. To mitigate such issues, this paper proposes three idioms based on aspect-oriented programming to implement flexible feature binding. We apply our three idioms, along with Edicts, to implement a flexible binding time for features in four different applications. By doing so, we were able to assess the resulting implementations by using software metrics that judge code-quality factors. Our evaluation suggests that our idioms reduce the above-mentioned problems when implementing flexible feature binding for the selected features. Rodrigo Andrade, Márcio Ribeiro 0001, Henrique Rebêlo, Paulo Borba, Vaidas Gasiunas, Lucas Satabin |
Comput. J. | 2 |
| 2016 | Assessing fine-grained feature dependencies
Iran Rodrigues, Márcio Ribeiro 0001, Flávio Medeiros, Paulo Borba, Baldoino Fonseca dos Santos Neto, Rohit Gheyi |
Inf. Softw. Technol. | 2 |
| 2015 | The Love/Hate Relationship with the C Preprocessor: An Interview StudyabstractThe C preprocessor has received strong criticism in academia, among others regarding separation of concerns, error proneness, and code obfuscation, but is widely used in practice. Many (mostly academic) alternatives to the preprocessor exist, but have not been adopted in practice. Since developers continue to use the preprocessor despite all criticism and research, we ask how practitioners perceive the C preprocessor. We performed interviews with 40 developers, used grounded theory to analyze the data, and cross-validated the results with data from a survey among 202 developers, repository mining, and results from previous studies. In particular, we investigated four research questions related to why the preprocessor is still widely used in practice, common problems, alternatives, and the impact of undisciplined annotations. Our study shows that developers are aware of the criticism the C preprocessor receives, but use it nonetheless, mainly for portability and variability. Many developers indicate that they regularly face preprocessor-related problems and preprocessor-related bugs. The majority of our interviewees do not see any current C-native technologies that can entirely replace the C preprocessor. However, developers tend to mitigate problems with guidelines, even though those guidelines are not enforced consistently. We report the key insights gained from our study and discuss implications for practitioners and researchers on how to better use the C preprocessor to minimize its negative impact. Flávio Medeiros, Christian Kästner, Márcio Ribeiro 0001, Sarah Nadi, Rohit Gheyi |
ECOOP | 3 |
| 2015 | An empirical study on configuration-related issues: investigating undeclared and unused identifiersabstractThe variability of configurable systems may lead to configuration-related issues (i.e., faults and warnings) that appear only when we select certain configuration options. Previous studies found that issues related to configurability are harder to detect than issues that appear in all configurations, because variability increases the complexity. However, little effort has been put into understanding configuration-related faults (e.g., undeclared functions and variables) and warnings (e.g., unused functions and variables). To better understand the peculiarities of configuration-related undeclared/unused variables and functions, in this paper we perform an empirical study of 15 systems to answer research questions related to how developers introduce these issues, the number of configuration options involved, and the time that these issues remain in source files. To make the analysis of several projects feasible, we propose a strategy that minimizes the initial setup problems of variability-aware tools. We detect and confirm 2 undeclared variables, 14 undeclared functions, 16 unused variables, and 7 unused functions related to configurability. We submit 30 patches to fix issues not fixed by developers. Our findings support the effectiveness of sampling (i.e., analysis of only a subset of valid configurations) because most issues involve two or less configuration options. Nevertheless, by analyzing the version history of the projects, we observe that a number of issues remain in the code for several years. Furthermore, the corpus of undeclared/unused variables and functions gathered is a valuable source to study these issues, compare sampling algorithms, and test and improve variability-aware tools. Flávio Medeiros, Iran Rodrigues, Márcio Ribeiro 0001, Leopoldo Teixeira, Rohit Gheyi |
GPCE | 3 |
| 2015 | Experience report: Evaluating the effectiveness of decision trees for detecting code smellsabstractDevelopers continuously maintain software systems to adapt to new requirements and to fix bugs. Due to the complexity of maintenance tasks and the time-to-market, developers make poor implementation choices, also known as code smells. Studies indicate that code smells hinder comprehensibility, and possibly increase change- and fault-proneness. Therefore, they must be identified to enable the application of corrections. The challenge is that the inaccurate definitions of code smells make developers disagree whether a piece of code is a smell or not, consequently, making difficult creation of a universal detection solution able to recognize smells in different software projects. Several works have been proposed to identify code smells but they still report inaccurate results and use techniques that do not present to developers a comprehensive explanation how these results have been obtained. In this experimental report we study the effectiveness of the Decision Tree algorithm to recognize code smells. For this, it was applied in a dataset containing 4 open source projects and the results were compared with the manual oracle, with existing detection approaches and with other machine learning algorithms. The results showed that the approach was able to effectively learn rules for the detection of the code smells studied. The results were even better when genetic algorithms are used to pre-select the metrics to use. Lucas Amorim, Evandro de Barros Costa, Nuno Antunes, Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001 |
ISSRE | 5 |
| 2015 | Ontology-based feature modeling: An empirical study in changing scenarios
Diego Dermeval, Thyago Tenório, Ig Ibert Bittencourt, Alan Silva, Seiji Isotani, Márcio Ribeiro 0001 |
Expert Syst. Appl. | 6 |
| 2015 | AutoRefactoring: A platform to build refactoring agents
Baldoino Fonseca dos Santos Neto, Márcio Ribeiro 0001, Viviane Torres da Silva, Christiano Braga, Carlos José Pereira de Lucena, Evandro de Barros Costa |
Expert Syst. Appl. | 2 |
| 2014 | Feature maintenance with emergent interfacesabstractHidden code dependencies are responsible for many complications in maintenance tasks. With the introduction of variable features in configurable systems, dependencies may even cross feature boundaries, causing problems that are prone to be detected late. Many current implementation techniques for product lines lack proper interfaces, which could make such dependencies explicit. As alternative to changing the implementation approach, we provide a tool-based solution to support developers in recognizing and dealing with feature dependencies: emergent interfaces. Emergent interfaces are inferred on demand, based on feature-sensitive intraprocedural and interprocedural data-flow analysis. They emerge in the IDE and emulate modularity benefits not available in the host language. To evaluate the potential of emergent interfaces, we conducted and replicated a controlled experiment, and found, in the studied context, that emergent interfaces can improve performance of code change tasks by up to 3 times while also reducing the number of errors. Márcio Ribeiro 0001, Paulo Borba, Christian Kästner |
ICSE | 1 |
| 2014 | Scaling Testing of Refactoring EnginesabstractProving refactoring sound with respect to a formal semantics is considered a challenge. In practice, developers write test cases to check their refactoring implementations. However, it is difficult and time consuming to have a good test suite since it requires complex inputs (programs) and an oracle to check whether it is possible to apply the transformation. If it is possible, the resulting program must preserve the observable behavior. There are some automated techniques for testing refactoring engines. Nevertheless, they may have limitations related to the program generator (exhaustiveness, setup, expressiveness), automation (types of oracles, bug categorization), time consumption or kinds of refactorings that can be tested. In this paper, we extend our previous technique to test refactoring engines. We improve expressiveness of the program generator for testing more kinds of refactorings, such as Extract Function. Moreover, developers just need to specify the input's structure in a declarative language. They may also set the technique to skip some consecutive test inputs to improve performance. We evaluate our technique in 18 refactoring implementations of Java (Eclipse and JRRT) and C (Eclipse). We identify 76 bugs (53 new bugs) related to compilation errors, behavioral changes, and overly strong conditions. We also compare the impact of the skip on the time consumption and bug detection in our technique. By using a skip of 25 in the program generator, it reduces in 96% the time to test the refactoring implementations while missing only 3.9% of the bugs. In a few seconds, it finds the first failure related to compilation error or behavioral change. Melina Mongiovi, Gustavo Mendes, Rohit Gheyi, Gustavo Soares, Márcio Ribeiro 0001 |
ICSME | 5 |
| 2013 | Investigating preprocessor-based syntax errorsabstractThe C preprocessor is commonly used to implement variability in program families. Despite the widespread usage, some studies indicate that the C preprocessor makes variability implementation difficult and error-prone. However, we still lack studies to investigate preprocessor-based syntax errors and quantify to what extent they occur in practice. In this paper, we define a technique based on a variability-aware parser to find syntax errors in releases and commits of program families. To investigate these errors, we perform an empirical study where we use our technique in 41 program family releases, and more than 51 thousand commits of 8 program families. We find 7 and 20 syntax errors in releases and commits of program families, respectively. They are related not only to incomplete annotations, but also to complete ones. We submit 8 patches to fix errors that developers have not fixed yet, and they accept 75% of them. Our results reveal that the time developers need to fix the errors varies from days to years in family repositories. We detect errors even in releases of well-known and widely used program families, such as Bash, CVS and Vim. We also classify the syntax errors into 6 different categories. This classification may guide developers to avoid them during development. Flávio Medeiros, Márcio Ribeiro 0001, Rohit Gheyi |
GPCE | 2 |
| 2013 | SPLLIFT: statically analyzing software product lines in minutes instead of yearsabstractA software product line (SPL) encodes a potentially large variety of software products as variants of some common code base. Up until now, re-using traditional static analyses for SPLs was virtually intractable, as it required programmers to generate and analyze all products individually. In this work, however, we show how an important class of existing inter-procedural static analyses can be transparently lifted to SPLs. Without requiring programmers to change a single line of code, our approach SPLLIFT automatically converts any analysis formulated for traditional programs within the popular IFDS framework for inter-procedural, finite, distributive, subset problems to an SPL-aware analysis formulated in the IDE framework, a well-known extension to IFDS. Using a full implementation based on Heros, Soot, CIDE and JavaBDD, we show that with SPLLIFT one can reuse IFDS-based analyses without changing a single line of code. Through experiments using three static analyses applied to four Java-based product lines, we were able to show that our approach produces correct results and outperforms the traditional approach by several orders of magnitude. Eric Bodden, Társis Tolêdo, Márcio Ribeiro 0001, Claus Brabrand, Paulo Borba, Mira Mezini |
PLDI | 3 |
| 2013 | Quantifying the effects of Aspectual Decompositions on Design by Contract Modularization: a Maintenance StudyabstractAlthough it is assumed that the implementation of design by contract is better modularized by means of aspect-oriented (AO) programming, there is no empirical evidence on the effectiveness of AO for modularizing non-trivial design by contract code in realistic development scenarios. This paper reports a quantitative and qualitative case study that evolves a real-life application to assess various facets of the adequacy of aspects for modularizing the design by contract concern. Our evaluation focused upon a number of system changes that are typically performed during software maintenance tasks. The study was driven by an analysis of fundamental modularity attributes, such as separation of concerns, coupling, conciseness, and change propagation. We have found that AO techniques improved separation of concerns and the design stability between the design by contract code and base application code throughout the development scenarios. However, contradicting the general intuition, the AO versions of the system did not present significant gains regarding four classical size metrics we employed. Henrique Rebêlo, Ricardo Massa Ferreira Lima, Uirá Kulesza, Márcio Ribeiro 0001, Yuanfang Cai, Roberta Coelho, Cláudio Sant'Anna, Alexandre Mota 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2013 | A design rule language for aspect-oriented programming
Alberto Costa Neto, Rodrigo Bonifácio, Márcio Ribeiro 0001, Carlos Eduardo Pontual, Paulo Borba, Fernando Castor Filho |
J. Syst. Softw. | 3 |
| 2012 | Enforcing Contracts for Aspect-oriented programs with Annotations, Pointcuts and Advice
Henrique Rebêlo, Ricardo Massa Ferreira Lima, Alexandre Mota 0001, César A. L. de Oliveira, Márcio Ribeiro 0001 |
SEKE | 5 |
| 2012 | Checking Contracts for AOP using XPIDRs
Henrique Rebêlo, Ricardo Massa Ferreira Lima, Alexandre Mota 0001, César A. L. de Oliveira, Márcio Ribeiro 0001 |
SEKE | 5 |
| 2011 | On the impact of feature dependencies when maintaining preprocessor-based software product linesabstractDuring Software Product Line (SPL) maintenance tasks, Virtual Separation of Concerns (VSoC) allows the programmer to focus on one feature and hide the others. However, since features depend on each other through variables and control-flow, feature modularization is compromised since the maintenance of one feature may break another. In this context, emergent interfaces can capture dependencies between the feature we are maintaining and the others, making developers aware of dependencies. To better understand the impact of code level feature dependencies during SPL maintenance, we have investigated the following two questions: how often methods with preprocessor directives contain feature dependencies? How feature dependencies impact maintenance effort when using VSoC and emergent interfaces? Answering the former is important for assessing how often we may face feature dependency problems. Answering the latter is important to better understand to what extent emergent interfaces complement VSoC during maintenance tasks. To answer them, we analyze 43 SPLs of different domains, size, and languages. The data we collect from them complement previous work on preprocessor usage. They reveal that the feature dependencies we consider in this paper are reasonably common in practice; and that emergent interfaces can reduce maintenance effort during the SPL maintenance tasks we regard here. Márcio Ribeiro 0001, Felipe Queiroz, Paulo Borba, Társis Tolêdo, Claus Brabrand, Sérgio Soares |
GPCE | 1 |
| 2011 | Assessing the Impact of Aspects on Design By Contract Effort: A Quantitative Study
Henrique Rebêlo, Ricardo Massa Ferreira Lima, Uirá Kulesza, Cláudio Sant'Anna, Roberta Coelho, Alexandre Mota 0001, Márcio Ribeiro 0001, César A. L. de Oliveira |
SEKE | 7 |
| 2008 | Developing Enterprise Applications with Support to Dynamic Unanticipated Evolution
Hyggo Oliveira de Almeida, Marcos F. Pereira, Márcio Ribeiro 0001, Angelo Perkusich, Emerson Loureiro, Evandro de Barros Costa |
SEKE | 3 |