Everton Guimarães

dblp:262/0667 · also Everton T. Guimarães · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-6740-6561ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Software Testing Education in the LLM Era: Insights and Emerging Theory
abstract
Large Language Models (LLMs) have been rapidly adopted across the software development lifecycle, becoming integral to modern engineering workflows. However, their use remains largely exploratory, with limited empirical studies and ongoing questions about their reliability. Despite this uncertainty, academic curricula must now adapt to reflect evolving industry practices. This raises a dilemma: how can educators integrate tools that remain under investigation yet are already essential for the next generation of professionals? This study takes early steps toward an empirically grounded theory of how LLMs are being integrated into software-testing education. Using an exploratory grounded design, we combine a focused review of five representative studies with a classroom case study in which graduate students contrasted manual and AI-assisted test-case design and reflected on their experiences. Developing such a grounded theoretical perspective provides a structured way to explain emerging educational dynamics and to guide future empirical and pedagogical work on responsible LLM integration.
Samhitha Dwarakanath, Nathalia Moraes do Nascimento, Everton Guimarães
AST3
2026 Empirical benchmarking of large language models for data science coding: a multidimensional evaluation
abstract
Abstract Despite growing enthusiasm for large language models (LLMs) as coding assistants, there remains limited empirical evidence of their effectiveness in domain-specific contexts such as data science. Existing benchmarks primarily focus on general-purpose programming and do not fully capture the challenges of data science tasks, which require data manipulation, statistical reasoning, algorithmic problem solving, and visualization. They also rarely assess practical dimensions such as first-attempt reliability, output consistency, error recovery, and cost efficiency. To address this gap, we introduce the LLM4DS-Benchmark and conduct a multidimensional empirical evaluation of seven LLMs—Gemini 2.5 Pro, Claude Sonnet 4.5, o3-mini, GPT-4.1, GPT-4o, Qwen3-Coder, and Perplexity Sonar—on 814 Python data science coding problems from StrataScratch platform, spanning Analytical, Algorithm, and Visualization tasks across three difficulty levels. Each problem received up to three attempts under a branching protocol that separates independent attempts from feedback-guided retries, enabling analysis of correctness, Pass@1, retry recovery, output consistency, execution behavior, visualization quality, code similarity, token usage, and cost per solved problem. Results show that Gemini 2.5 Pro achieved the highest overall success rate (81.3%) and Pass@1 (62.2%), but at a median cost of \$0.10740 per solved problem—316 times higher than Qwen3-Coder (\$0.00034). Across models, retries improved performance by 19–24 percentage points, with feedback resolving 20–31% of initial failures. Output consistency varied significantly across identical prompts, particularly for Analytical tasks. Model rankings also shifted by task type and evaluation dimension, with no single model dominating across all categories. Instead, a Pareto-optimal set—Qwen3-Coder, GPT-4.1, o3-mini, Claude Sonnet 4.5, and Gemini 2.5 Pro—emerged, reflecting trade-offs among accuracy, cost, and reliability. These findings highlight the need for multidimensional, task-aware benchmarking and suggest that model selection for data science coding should be guided by task characteristics and practical constraints rather than aggregate success rate alone.
Santhosh Anitha Boominathan, Sai Sanjna Chintakunta, Everton Guimarães, Kailasam Satyamurthy, Nathalia Moraes do Nascimento
Empir. Softw. Eng.3
2025 Analyzing Prominent LLMs: An Empirical Study of Performance and Complexity in Solving LeetCode Problems
abstract
The rapid advancement of Generative AIs (GenAIs), particularly Large Language Models (LLMs), has transformed software engineering by automating tasks such as code generation, testing, and debugging. As these models become increasingly integrated into development workflows, evaluating their performance systematically is crucial for optimizing their effectiveness in real-world applications. This study aims to benchmark six prominent LLMs—ChatGPT, Copilot, Gemini, Claude, Perplexity, and DeepSeek—on algorithm and data structure problems from LeetCode, assessing their strengths and limitations in solving programming challenges. The study evaluates LLM performance on 150 LeetCode problems, generating solutions in Java and Python. Performance metrics include execution time, memory usage, and computational complexity (time and space). A joint analysis ranks the models based on multiple performance factors. The evaluation reveals variations in LLM efficiency, with some models consistently outperforming others across different difficulty levels. Copilot, DeepSeek, and Perplexity demonstrate strong performance, while Gemini struggles with harder problems. Differences in execution time and memory usage are also noted across programming languages. The findings contribute to a deeper understanding of LLM capabilities in code generation and provide insights to help developers make informed decisions when selecting LLMs based on problem complexity and programming language.
Everton Guimarães, Nathalia Moraes do Nascimento, Asish Nelapati, Chandan Shivalingaiah
EASE1
2025 How Effective are LLMs for Data Science Coding? A Controlled Experiment
abstract
The adoption of Large Language Models (LLMs) for code generation in data science offers substantial potential for enhancing tasks such as data manipulation, statistical analysis, and visualization. However, the effectiveness of these models in the data science domain remains underexplored. This paper presents a controlled experiment that empirically assesses the performance of four leading LLM-based AI assistants-Microsoft Copilot (GPT-4 Turbo), ChatGPT (o1-preview), Claude (3.5 Sonnet), and Perplexity Labs (Llama-3.1-70b-instruct)-on a diverse set of data science coding challenges sourced from the Stratacratch platform. Using the Goal-Question-Metric (GQM) approach, we evaluated each model’s effectiveness across task types (Analytical, Algorithm, Visualization) and varying difficulty levels. Our statistical testing confirms that all models achieved success rates significantly above $50 \%$, demonstrating performance beyond chance. ChatGPT and Claude significantly exceeded the $60 \%$ threshold, but no model reached $70 \%$, indicating limitations in achieving higher accuracy. ChatGPT maintained consistent performance across difficulty levels, whereas Claude’s success varied with task complexity. Hypothesis testing indicates that task type does not significantly impact success rate overall. For analytical tasks, efficiency analysis shows no significant differences in execution times, though ChatGPT tended to be slower and less predictable despite high success rates. For visualization tasks, while similarity quality among LLMs is comparable, ChatGPT consistently delivered the most accurate outputs. This study provides a structured, empirical evaluation of LLMs in data science, delivering insights that support informed model selection tailored to specific task demands. Our findings establish a framework for future AI assessments, emphasizing the value of rigorous evaluation beyond basic accuracy measures.
Nathalia Moraes do Nascimento, Everton Guimarães, Sai Sanjna Chintakunta, Santhosh Anitha Boominathan
MSR2
2023 Managing Technical Debt Using Intelligent Techniques - A Systematic Mapping Study
abstract
Technical Debt (TD) is a metaphor reflecting technical compromises that can yield short-term benefits but might hurt the long-term health of a software system. With the increasing amount of data generated when performing software development activities, an emergent research field has gained attention: applying Intelligent Techniques to solve Software Engineering problems. Intelligent Techniques were used to explore data for knowledge discovery, reasoning, learning, planning, perception, or supporting decision-making. Although these techniques can be promising, there is no structured understanding related to their application to support Technical Debt Management (TDM) activities. Within this context, this study aims to investigate to what extent the literature has proposed and evaluated solutions based on Intelligent Techniques to support TDM activities. To this end, we performed a Systematic Mapping Study (SMS) to investigate to what extent the literature has proposed and evaluated solutions based on Intelligent Techniques to support TDM activities. In total, 150 primary studies were identified and analyzed, dated from 2012 to 2021. The results indicated a growing interest in applying Intelligent Techniques to support TDM activities, the most used: Machine Learning and Reasoning under uncertainty. Intelligent Techniques aimed to assist mainly TDM activities related to identification, measurement, and monitoring. Design TD, Code TD, and Architectural TD are the TD types in the spotlight. Most studies were categorized at automation levels 1 and 2, meaning that existing approaches still require substantial human intervention. Symbolists and Analogizers are levels of explanation presented by most Intelligent Techniques, implying that these solutions conclude a general truth after considering a sufficient number of particular cases. Moreover, we also cataloged the empirical research types, contributions, and validation strategies described in primary studies. Based on our findings, we argue that there is still room to improve the use of Intelligent Techniques to support TDM activities. The open issues that emerged from this study can represent future opportunities for practitioners and researchers.
Danyllo Albuquerque, Everton Guimarães, Graziela Tonin, Pilar Rodríguez 0002, Mirko Barbosa Perkusich, Hyggo Oliveira de Almeida, Angelo Perkusich, Ferdinandy Chagas
IEEE Trans. Software Eng.2
2022 Comprehending the use of intelligent techniques to support technical debt management
abstract
Technical Debt (TD) refers to the consequences of taking shortcuts when developing software. Technical Debt Management (TDM) becomes complex since it relies on a decision process based on multiple and heterogeneous data, which are not straightforward to be synthesized. In this context, there is a promising opportunity to use Intelligent Techniques to support TDM activities since these techniques explore data for knowledge discovery, reasoning, learning, or supporting decision-making. Although these techniques can be used for improving TDM activities, there is no empirical study exploring this research area. This study aims to identify and analyze solutions based on Intelligent Techniques employed to support TDM activities. A Systematic Mapping Study was performed, covering publications between 2010 and 2020. From 2276 extracted studies, we selected 111 unique studies. We found a positive trend in applying Intelligent Techniques to support TDM activities, being Machine Learning, Reasoning Under Uncertainty, and Natural Language Processing the most recurrent ones. Identification, measurement, and monitoring were the more recurrent TDM activities, whereas Design, Code, and Architectural were the most frequently investigated TD types. Although the research area is up-and-coming, it is still in its infancy, and this study provides a baseline for future research.
Danyllo Albuquerque, Everton Guimarães, Graziela Tonin, Mirko Barbosa Perkusich, Hyggo Oliveira de Almeida, Angelo Perkusich
TechDebt@ICSE2
2021 A Comparative Study of Psychometric Instrumentsin Software Engineering
abstract
Over the years, researchers have explored the influence of human factors in software engineering, showing that the team members' personalities might affect teamwork.However, it is challenging to measure software engineers' personalities due to the number of available psychometric instruments and the possibility of using different scales and classifications.Our study compares the personality traits measured by three psychometric instruments used in Software Engineering: Big Five Inventory (BFI), 16 Personality Factors (16PF), and Context Cards (CC).For this purpose, we executed an empirical study in which we collected data from 29 software developers for each of the evaluated instruments.As a result, we identified a moderate correlation between BFI and 16PF, confirming the current stateof-the-art.For the remaining combinations, there was a weak correlation.As implications for this research, there is a need to empirically evaluate BFI and CC (context-specific survey) in terms of construct validity since they have moderate to low correlation.
Gleyser Guimarães, Mirko Barbosa Perkusich, Danyllo Albuquerque, Everton Guimarães, Danilo Santos 0001, Hyggo Oliveira de Almeida, Angelo Perkusich
SEKE4
2018 On the UML use in the Brazilian industry: A state of the practice survey (S)
abstract
Context: The Unified Modeling Language (UML) has become the standard for modeling software.Several surveys on the UML usage have been proposed in recent years.However, none of them explores the UML use in specific regional scope, and thus little is known about the practices and perceptions of UML use from the perspective of practitioners in the Brazilian industry.Objective: This paper reports on a survey focused on identifying the state-of-the-practice of the Brazilian industry for what concerns the UML usage in real-world settings.Method: In total, 222 practitioners from 140 different Information Technology companies have answered an on-line (or printed) questionnaire concerning their UML use experiences, the difficulty in adopting UML and what should be done to increase the UML adoption in practice.Result: The results show that: (1) 60 participants (28.2%) have used UML in their daily work, while 73.2% have not; (2) 55.41% of the surveyed participants did not disagree with the statement that UML is the "lingua franca" in software modeling; (3) 61.26% reported to find that the automatic creation of UML diagrams to represent a big picture of the system under development would be useful to boost UML use.Conclusion: The UML is not often used in the work life of participants.In addition, no relationship was identified between the use of UML and the participant company being a software factory.
Kleinner Farias, Lucian Gonçales, Vinicius Bischoff, Bruno da Silva 0002, Everton Guimarães, Jacob Nogle
SEKE5
2018 Effects of Model Composition Techniques on Effort and Affective States: A Controlled Experiment (S)
abstract
Even though existing heuristics and specification-based techniques support composing design models, it is still considered a time-consuming and highly intensive task.In addition, there is a lack of studies exploring the effects of composition techniques on software developers' affective state and development effort.This study reports a pilot study to investigate these effects while developers apply composition techniques to detect and resolve inconsistencies in output-composed models.In this sense, a widely known wearable EEG headset, namely Emotiv EPOC, with 14 channels was used, while developers made use of heuristic-based and specification-based composition techniques to evolve design models.Our results suggest that using heuristic-based techniques produced a higher effect on the developers' affectivity, compared to specification-based techniques.Moreover, the higher the effects on the developers' affectivity, the higher the odds to invest less effort and produce correctly composed design models.
Mateus Manica, Kleinner Farias, Lucian Gonçales, Vinicius Bischoff, Bruno da Silva 0002, Everton Guimarães
SEKE6
2018 Exploring architecture blueprints for prioritizing critical code anomalies: Experiences and tool support
abstract
Summary The manifestation of code anomalies in software systems often indicates symptoms of architecture degradation. Several approaches have been proposed to detect such anomalies in the source code. However, most of them fail to assist developers in prioritizing anomalies harmful to the software architecture of a system. This article presents an investigation on how developers, when supported by architecture blueprints, are able to prioritize architecturally relevant code anomalies. First, we performed a controlled experiment where participants explored both blueprints and source code to reveal architecturally relevant code anomalies. Although the use of blueprints has the potential to improve code anomaly prioritization, the participants often made several mistakes. We found these mistakes might occur because developers miss relationships between implementation and blueprint elements when they prioritize anomalies in an ad hoc manner. Furthermore, the time spent on the prioritization process was considerably high. Aiming to improve the accuracy and effectiveness of the process, we provided means to automate the prioritization process. In particular, we explored 3 prioritization criteria, which establish different ways of relating the blueprint elements with code anomalies. These criteria were implemented in the JSpIRIT tool. The approach was evaluated in the context of 2 applications with satisfactory precision results.
Everton Guimarães, Santiago A. Vidal, Alessandro F. Garcia 0001, Jorge Andrés Díaz Pace, Claudia A. Marcos
Softw. Pract. Exp.1
2015 JSAN: A Framework to Implement Normative Agents
abstract
Norms have become a promising mechanism to ensure that open multi-agent systems (MASs) produce a desirable social outcome.MASs can be defined as societies in which autonomous agents work to achieve both societal and individual goals.Norms regulate the behavior of agents by defining permissions, obligations and prohibitions, as well as encouraging and discouraging the fulfillment of norms through rewards and punishments mechanisms.Once the priority of software agent is the satisfaction of its own desires and goals, each agent must evaluate the effects associated to the fulfillment or violation of one or more norms before choosing which one should be complied.This paper introduces a framework for normative MASs simulation that provides mechanisms for understanding the impact of norms on an agent and the society to which an agent belongs.
Marx L. Viana, Paulo S. C. Alencar, Everton Guimarães, Francisco J. P. Cunha, Donald D. Cowan, Carlos José Pereira de Lucena
SEKE3
2014 Exploring Blueprints on the Prioritization of Architecturally Relevant Code Anomalies - A Controlled Experiment
abstract
The progressive insertion of code anomalies in evolving programs may lead to architecture degradation symptoms. Several approaches have been proposed aiming to detect code anomalies in the source code, such as God Class and Shotgun Surgery. However, most of them fail to assist developers on prioritizing code anomalies harmful to the software architecture. These approaches often rely on source code analysis and do not provide developers with useful information to help the prioritization of those anomalies that impact on the architectural design. In this context, this paper presents a controlled experiment aiming at investigating how developers, when supported by architecture blueprints, are able to prioritize different types of code anomalies in terms of their architectural relevance. Our contributions include: (i) quantitative indicators on how the use of blueprints may improve process of prioritizing code anomalies, (ii) a discussion of how blueprints may help on the prioritization processes, (iii) an analysis of whether and to what extent the use of blueprints impacts on the time for revealing architecturally relevant code anomalies, and (iv) a discussion on the main characteristics of false positives and false negatives observed by the actual developers.
Everton Guimarães, Alessandro F. Garcia 0001, Yuanfang Cai
COMPSAC1
2013 Prioritizing software anomalies with software metrics and architecture blueprints: a controlled experiment
abstract
According to recent studies, architecture degradation is to a large extent a consequence of the introduction of code anomalies as the system evolves. Many approaches have been proposed for detecting code anomalies, but none of them has been efficient on prioritizing code anomalies that represent real problems in the architecture design. In this sense, our work aims to investigate whether the prioritization of instances of three types of classical code anomalies, Divergent Change, God Class and Shotgun Surgery, can be improved when supported by architecture blueprints. These blueprints are informal models often available in software projects, and they are used to capture key architecture decisions. Moreover, we are also investigating what information may be useful in the design blueprints to help developers on prioritizing the most critical software anomalies. In many cases, developers indicated that it would be interesting the insertion of additional information on the blueprints in order to detect architecturally-relevant anomalies.
Everton Guimarães, Alessandro F. Garcia 0001, Eduardo Figueiredo 0001, Yuanfang Cai
MiSE1