João Brunet

dblp:41/4431 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Benchmark Data Contamination in Underrepresented Languages: A Comprehensive Analysis Using Brazilian Data
Iriedson Souto Maior de Moraes Vilar, David Candeia Maia, João Brunet, Fábio Morais 0001, Leandro Balby Marinho
LREC3
2026 Lexical and semantic representations for similar bug report detection: A TF-IDF and frozen T5 comparative study
abstract
Abstract In software development, bug reports (BRs) are essential for identifying defects, but the volume of reports in large projects makes manual relatedness analysis slow and error-prone. We study machine-learning approaches for predicting BR relatedness under a file-overlap target, File-Change Similarity (FCS). We compare TF-IDF, frozen sentence-level T5 embeddings (without domain fine-tuning), and a hybrid lexical-semantic representation. Our pipeline covers data retrieval, preprocessing, vectorization, normalization, neural-network training, and evaluation. We evaluated 56 models using various modeling strategies. Analysis reveals that using complete vectors as features is more effective than cosine distance. The hybrid approach shows competitive descriptive performance comparable to TF-IDF alone. Fine-tuning on 14 models tested 168 hyperparameter combinations, with Adam and RMSprop optimizers showing best performance. Key contributions include evaluating T5 and TF-IDF performance for BRs, exploring a hybrid approach, and providing a methodological framework for representation comparison. This research offers suggestions for improving efficiency in development and resource allocation. In the context of frozen T5 embeddings, the findings on T5 performance and the comparison with strong TF-IDF baselines drive future research directions. Since the T5 weights were not specifically trained on the bug report domain (frozen), these results serve as a baseline for future fine-tuning experiments.
Iann Barbosa, João Brunet, Franklin Ramalho
Softw. Qual. J.2
2025 Beyond Functionality: Automating Algorithm Design Evaluation in Introductory Programming Courses
Caio Oliveira, Leandra Silva, João Brunet
CSEDU (2)3
2025 A Defect Taxonomy for Infrastructure as Code: A Replication Study
abstract
Background: As Infrastructure as Code (IaC) becomes standard practice, ensuring the reliability of IaC scripts is essential. Defect taxonomies are valuable tools for this, offering a common language for issues and enabling systematic tracking. A significant prior study developed such a taxonomy, but based it exclusively on the declarative language Puppet. It remained unknown whether this taxonomy applies to programming language-based IaC (PLIaC) tools like Pulumi, Terraform CDK, and AWS CDK. Aim: We replicated this foundational work to assess the generalizability of the taxonomy across a broader and more diverse landscape. Method: We performed qualitative analysis on 3,364 defect-related commits from 285 open-source PL-IaC repositories (PIPr dataset) to derive a PL-IaC specific defect taxonomy. We then enhanced the ACID tool, originally developed for the prior study, to automatically classify and analyze defect distributions across an expanded dataset- 447 open-source repositories and 94 proprietary projects from VTEX (e-commerce) and Nubank (financial)-incorporating modern PL-IaC tools absent in the original work. Results: Our research confirmed the same eight defect categories identified in the original study, with idempotency and security defects appearing infrequently but persistently across projects. Configuration Data defects maintain high frequency in both open-source and proprietary codebases. Despite differences in project types, overall defect proportions remain similar. Conclusions: Our replication supports the generalizability of the original taxonomy, suggesting IaC development challenges surpass organizational boundaries. Configuration Data defects emerge as a persistent highfrequency problem, while idempotency and security defects remain important concerns despite lower frequency. These patterns appear consistent across open-source and proprietary projects, indicating they are fundamental to the IaC paradigm itself, transcending specific tools or project types.
Wendell Oliveira, Filipe Paiva, Thiago Emmanuel Pereira, João Brunet
ESEM4
2022 A Systematic Literature Review on Predictive Cognitive Skills in Novice Programming
abstract
This Research Full Paper presents a Systematic Literature Review (SLR) investigating predictive programming skills and strategies to foster and measure such skills. Predictive skills are specific skills that precede a milestone or the development of other more structured skills, in this case, skills that precede programming skills. Introductory Programming Courses (CS1) feature students with a wide range of skill levels. This difference causes educators to get lost when formulating their practices to teach such a diverse group. There is no clear vision about what previous skills professors should foster before a student enters a CS1, much less which strategies to foster/measure such skills. Due to this limitation, we present/plan an RSL following the guidelines proposed by Kitchenham (2004) to achieve these goals. Our main results are a) Predictive programming skills are problem-solving, abstract thinking, mathematical reasoning, and cognitive flexibility; b) Different researchers use different approaches to teach programming skills based on different educational theories, teaching frameworks, or educational approaches; c) Several studies use the Classical Test Theory as a way to measure predictive programming skills. However, some universities have adopted other theories for this practice, such as Item Response Theory.
Jucelio S. Santos, Wilkerson de L. Andrade, João Brunet, Monilly Ramos Araujo de Melo
FIE3
2020 Applying Item Response Theory to Evaluate Instruments of Introductory Programming Skills Measurement
abstract
This Research-to-practice Full Paper presents an exploratory and preliminary investigation on the reliability and validity of instruments for measuring introductory programming skills. Our data set consists of the performance of 30 students who participated in the experiment of a Brazilian university. We provide participants with instructional material, practical problems and solutions based on different experimental conditions. The results suggest that the instruments have a good internal consistency index and items with excellent psychometric properties. In addition, some initial evidence to suggest practicing each of the four programming skills. We found that participants who received practice in all four skills obtained a better estimate of study material and Assessment, especially for the most advanced knowledge of programming introduction.
Jucelio S. Santos, Wilkerson de L. Andrade, João Brunet, Monilly Ramos Araujo de Melo
FIE3
2020 A Systematic Literature Review of Methodology of Learning Evaluation Based on Item Response Theory in the Context of Programming Teaching
abstract
This Research Full Paper presents a Systematic Literature Review (SLR) that investigates state-of-the-art methods, processes, approaches, and instruments based on Item Response Theory (IRT) in the programming teaching context. Various studies support professors and students to improve the evaluation process in teaching programming. Among the different techniques and theories associated with these studies, the IRT gained prominence because of its effectiveness. Due to the lack of an overview in the area, we present a SLR with the objective of identity studies that use IRT as a methodology of evaluation of learning in the context of Programming Teaching; studies which present real-time feedback; and experimental studies that have been carried out to validate them. To achieve these goals, we planned an SLR following the guidelines proposed by Kitchenham (2004). Our main findings are a) There is a limited number of studies that explore IRT in evaluation methodologies in the teaching programming context. Among these studies, the main focus was the development on of instruments to measure programming skills; b) Most instruments are multiple-choice, adopt the 3PL model, and do not show evidence of real-time feedback except SIETTE and CodeWorkout and are also the only instruments that evaluate the coding capacity of individuals; c) The studies showed scientific evidence on the use of IRT in the evaluation methodology in the programming teaching context. The studies presented experiments that indicate a significant gain in the measurement of programming abilities when compared to the traditional evaluation method.
Jucelio S. Santos, Wilkerson de L. Andrade, João Brunet, Monilly Ramos Araujo de Melo
FIE3
2018 Improving Tail Latency of Stateful Cloud Services via GC Control and Load Shedding
abstract
Most of the modern cloud web services execute on top of runtime environments like .NET's Common Language Runtime or Java Runtime Environment. On the one hand, runtime environments provide several off-the-shelf benefits like code security and cross-platform execution. On the other hand, runtime's features such as just-in-time compilation and automatic memory management add a non-deterministic overhead to the overall service time, increasing the tail of the latency distribution. In this context, the Garbage Collector (GC) is among the leading causes of high tail latency. To tackle this problem, we developed the Garbage Collector Control Interceptor (GCI) - a request interceptor algorithm, which is agnostic regarding the cloud service language, internals, and its incoming load. GCI is wholly decentralized and improves the tail latency of cloud services by making sure that service instances shed the incoming load while cleaning up the runtime heap. We evaluated GCI's effectiveness in a stateful service prototype, varying the number of available instances. Our results showed that using GCI eliminates the impact of the garbage collection on the service latency for small (4 nodes) and large (64 nodes) deployments with no throughput loss.
Daniel Fireman, João Brunet, Raquel Lopes 0001, David Quaresma, Thiago Emmanuel Pereira
CloudCom2
2018 Can students help themselves? An investigation of students' feedback on the quality of the source code
abstract
This Research to Practice Full Paper presents a study on the evaluation of qualitative aspects of students' programs in an introductory programming course. Approaches have been proposed in order to address the quality of the source code, but they typically focus on automated analysis of syntactic aspects which might lead to generic feedback. In this study, we investigate if, by including students as evaluators, we could provide personalized feedback on the quality of source code. To do so, we applied a survey with assignments and their respective source codes answered by students in previous terms. Teachers and students analyze those source codes and gave suggestions to improve them qualitatively and, we found that most students identified code quality aspects with a similarity equal to or greater than 50% in comparison to teachers' and that similarity increases as students progress in the course. We found that students are particularly good at finding and giving feedback on complexity issues. This study may lead to further investigations on addressing source code quality on collaborative learning, and may also support the development of lint-like tools, once it yields detailed information on how students provide feedback regarding source code quality.
Raul Andrade, João Brunet
FIE2
2018 Automatic Decomposition of Java Open Source Pull Requests: A Replication Study
Victor da C. Luna Freire, João Brunet, Jorge C. A. de Figueiredo
SOFSEM2
2015 Helping Developers Help Themselves: Automatic Decomposition of Code Review Changesets
abstract
Code Reviews, an important and popular mechanism for quality assurance, are often performed on a change set, a set of modified files that are meant to be committed to a source repository as an atomic action. Understanding a code review is more difficult when the change set consists of multiple, independent, code differences. We introduce CLUSTERCHANGES, an automatic technique for decomposing change sets and evaluate its effectiveness through both a quantitative analysis and a qualitative user study.
Michael Barnett 0001, Christian Bird, João Brunet, Shuvendu K. Lahiri
ICSE (1)3
2014 Do developers discuss design?
abstract
Design is often raised in the literature as important to attaining various properties and characteristics in a software system. At least for open-source projects, it can be hard to find evidence of ongoing design work in the technical artifacts produced as part of the development. Although developers usually do not produce specific design documents, they do communicate about design in different ways. In this paper, we provide quantitative evidence that developers address design through discussions in commits, issues, and pull requests. To achieve this, we built a discussions' classifier and automatically labeled 102,122 discussions from 77 projects. Based on this data, we make four observations about the projects: i) on average, 25% of the discussions in a project are about design; ii) on average, 26% of developers contribute to at least one design discussion; iii) only 1% of the developers contribute to more than 15% of the discussions in a project; and iv) these few developers who contribute to a broad range of design discussions are also the top committers in a project.
João Brunet, Gail C. Murphy, Ricardo Terra, Jorge C. A. de Figueiredo, Dalton Serey Guerrero
MSR1
2013 Measuring the Structural Similarity between Source Code Entities (S)
Ricardo Terra, João Brunet, Luis Fernando Miranda, Marco Túlio Valente, Dalton Serey Guerrero, Douglas Castilho 0001, Roberto da Silva Bigonha
SEKE2
2011 Structural conformance checking with design tests: An evaluation of usability and calability
abstract
Verifying whether a software meets its functional requirements plays an important role in software development. However, this activity is necessary, but not sufficient to assure software quality. It is also important to check whether the code meets its design specification. Although there exists substantial tool support to assure that a software does what it is supposed to do, verifying whether it conforms to its design remains as an almost completely manual activity. In a previous work, we proposed design tests — test-like programs that automatically check implementations against design rules. Design test is an application of the concept of test to design conformance checking. To support design tests for Java projects, we developed DesignWizard, an API that allows developers to write and execute design tests using the popular JUnit testing framework. In this work, we present a study on the usability and scalability of DesignWizard to support structural conformance checking through design tests. We conducted a qualitative usability evaluation of DesignWizard using the Think Aloud Protocol for APIs. In the experiment, we challenged eleven developers to compose design tests for an open-source software project. We observed that the API meets most developers' expectations and that they had no difficulties to code design rules as design tests. To assess its scalability, we evaluated DesignWizard's use of CPU time and memory consumption. The study indicates that both are linear functions of the size of software under verification.
João Brunet, Dalton Serey Guerrero, Jorge C. A. de Figueiredo
ICSM1