VLDB 2026 Research / reviewers in the wild / expert
Amanda Santana
dblp:276/3266
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-1969-3460ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating a continuous feedback strategy to enhance machine learning code smell detection
Daniel Cruz, Amanda Santana, Eduardo Figueiredo 0001 |
Sci. Comput. Program. | 2 |
| 2024 | Unraveling the Impact of Code Smell Agglomerations on Code StabilityabstractCode smells are symptoms in the source code that indicate code quality degradation and, consequently, may affect code comprehension and maintenance. Moreover, when two or more code smells occur on the same piece of code, forming an agglomeration, they may be more harmful to the code quality. Although the impact of smells in isolation is well known, the impact of their agglomeration is still underexplored. Our goal with this study is to provide evidence of how agglomerations impact code stability, i.e. we investigate if agglomeration suffers more modifications along the system evolution, and in which intensity. For this purpose, we mined two years of commit history from 30 open-source Java systems from GitHub. To analyze code stability, we considered four measurements: the number of commits, lines of modified code, rate of modified classes, and the proportion of changes. We examined these measurements from two perspectives: by system and by aggregating all system data. Additionally, we further considered how a class created/deleted in this time span impacts our results. Our main findings are: (i) classes with two or more code smells of different types change more frequently and in more intensity than classes with a single smell or no smell; (ii) the stability of the class varies greatly with the system under analysis; (iii) when a smelly class was deleted in our time range, they usually had several lines of code added until it became unsustainable. We can conclude that agglomerations change with more frequency and intensity, raising maintenance and evolution costs. Consequently, this information can be used to prioritize code refactoring. Amanda Santana, Eduardo Figueiredo 0001, Juliana Alves Pereira |
ICSME | 1 |
| 2024 | Tuning Code Smell Prediction Models: A Replication StudyabstractIdentifying code smells in projects is a non-trivial task, and it is often a subjective activity since developers have different understandings about them. The use of machine learning techniques to predict code smells is gaining attention. In this replication study, our goals are: (i) verify if previous model's performance maintain when we extract data from updated systems; and (ii) explore and provide evidences of how the use of different feature engineering and resampling techniques can enhance code smell prediction model's performance. For these purposes, we evaluate four smells: God Class, Refused Bequest, Feature Envy and Long Method. We first replicate a previous study that focus on the algorithm's performance to identify the best models for each smell using a different dataset composed of 30 Java systems. This first experiment provides us a baseline model that is used in the second experiment. In the second experiment, we compare the performance of the baseline model with other models tuned with polynomial features and resample techniques. Our main results are: for datasets with imbalances lower than a ratio of 1:100, such as God Class and Long Method, the use of oversample techniques yielded better results. For datasets with more severe imbalance, like Refused Bequest and Feature Envy, the undersample techniques performed better. The feature selection technique, despite a minor impact on the results, provided insights. For instance, we need new features to represent some code smells, such as Long Method and Feature Envy. Henrique Gomes Nunes, Amanda Santana, Eduardo Figueiredo 0001, Heitor A. X. Costa |
ICPC | 2 |
| 2024 | An exploratory evaluation of code smell agglomerations
Amanda Santana, Eduardo Figueiredo 0001, Juliana Alves Pereira, Alessandro F. Garcia 0001 |
Softw. Qual. J. | 1 |
| 2023 | Yet Another Model! A Study on Model's Similarities for Defect and Code SmellsabstractAbstract Software defect and code smell prediction help developers identify problems in the code and fix them before they degrade the quality or the user experience. The prediction of software defects and code smells is challenging, since it involves many factors inherent to the development process. Many studies propose machine learning models for defects and code smells. However, we have not found studies that explore and compare these machine learning models, nor that focus on the explainability of the models. This analysis allows us to verify which features and quality attributes influence software defects and code smells. Hence, developers can use this information to predict if a class may be faulty or smelly through the evaluation of a few features and quality attributes. In this study, we fill this gap by comparing machine learning models for predicting defects and seven code smells. We trained in a dataset composed of 19,024 classes and 70 software features that range from different quality attributes extracted from 14 Java open-source projects. We then ensemble five machine learning models and employed explainability concepts to explore the redundancies in the models using the top-10 software features and quality attributes that are known to contribute to the defects and code smell predictions. Furthermore, we conclude that although the quality attributes vary among the models, the complexity, documentation, and size are the most relevant. More specifically, Nesting Level Else-If is the only software feature relevant to all models. Geanderson E. dos Santos, Amanda Santana, Gustavo Vale, Eduardo Figueiredo 0001 |
FASE | 2 |
| 2020 | Detecting bad smells with machine learning algorithms: an empirical studyabstractBad smells are symptoms of bad design choices implemented on the source code. They are one of the key indicators of technical debts, specifically, design debt. To manage this kind of debt, it is important to be aware of bad smells and refactor them whenever possible. Therefore, several bad smell detection tools and techniques have been proposed over the years. These tools and techniques present different strategies to perform detections. More recently, machine learning algorithms have also been proposed to support bad smell detection. However, we lack empirical evidence on the accuracy and efficiency of these machine learning based techniques. In this paper, we present an evaluation of seven different machine learning algorithms on the task of detecting four types of bad smells. We also provide an analysis of the impact of software metrics for bad smell detection using a unified approach for interpreting the models' decisions. We found that with the right optimization, machine learning algorithms can achieve good performance (F1 score) for two bad smells: God Class (0.86) and Refused Parent Bequest (0.67). We also uncovered which metrics play fundamental roles for detecting each bad smell. Daniel Cruz, Amanda Santana, Eduardo Figueiredo 0001 |
TechDebt@ICSE | 2 |