Henrique Gomes Nunes

dblp:355/9953 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-0982-3014ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
abstract
Large Language Models (LLMs) have gained attention for addressing coding problems, but their effectiveness in fixing code maintainability remains unclear. This study evaluates LLMs capability to resolve 127 maintainability issues from 10 GitHub repositories. We use zero-shot prompting for Copilot Chat and Llama 3.1, and few-shot prompting with Llama only. The LLM-generated solutions are assessed for compilation errors, test failures, and new maintainability problems. Llama with few-shot prompting successfully fixed 44.9 % of the methods, while Copilot Chat and Llama zero-shot fixed 32.29 % and 30 %, respectively. However, most solutions introduced errors or new maintainability issues. We also conducted a human study with 45 participants to evaluate the readability of 51 LLM-generated solutions. The human study showed that 68.63 % of participants observed improved readability. Overall, while LLMs show potential for fixing maintainability issues, their introduction of errors highlights their current limitations.
Henrique Gomes Nunes, Eduardo Figueiredo 0001, Larissa Rocha Soares, Sarah Nadi, Fischer Ferreira, Geanderson E. dos Santos
SANER1
2024 Tuning Code Smell Prediction Models: A Replication Study
abstract
Identifying code smells in projects is a non-trivial task, and it is often a subjective activity since developers have different understandings about them. The use of machine learning techniques to predict code smells is gaining attention. In this replication study, our goals are: (i) verify if previous model's performance maintain when we extract data from updated systems; and (ii) explore and provide evidences of how the use of different feature engineering and resampling techniques can enhance code smell prediction model's performance. For these purposes, we evaluate four smells: God Class, Refused Bequest, Feature Envy and Long Method. We first replicate a previous study that focus on the algorithm's performance to identify the best models for each smell using a different dataset composed of 30 Java systems. This first experiment provides us a baseline model that is used in the second experiment. In the second experiment, we compare the performance of the baseline model with other models tuned with polynomial features and resample techniques. Our main results are: for datasets with imbalances lower than a ratio of 1:100, such as God Class and Long Method, the use of oversample techniques yielded better results. For datasets with more severe imbalance, like Refused Bequest and Feature Envy, the undersample techniques performed better. The feature selection technique, despite a minor impact on the results, provided insights. For instance, we need new features to represent some code smells, such as Long Method and Feature Envy.
Henrique Gomes Nunes, Amanda Santana, Eduardo Figueiredo 0001, Heitor A. X. Costa
ICPC1