Italo Uchoa

dblp:388/0378 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0002-0115-5461ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 OmniCCG: Agnostic Code Clone Genealogy Extractor
abstract
When two or more code snippets are identical or sufficiently similar, they form code clones. Such duplication can harm system maintainability as the software evolves. Code clone genealogy (CCG) extraction involves analyzing successive versions of a software system to identify code clones, their modifications, additions, and removals. Visualizing clone genealogies helps developers manage their clones, improving code comprehensibility and maintainability. Despite their importance, to the best of our knowledge, no fully functional, easily executable clone genealogy extractor exists. Furthermore, all extractors proposed in the literature are specifically designed to work with a particular set of clone detectors, resulting in strong coupling. To address these shortcomings, this paper presents OmniCCG, a code clone genealogy extractor that is agnostic to clone detectors. Given a Git repository and user settings, OmniCCG extracts code clone genealogies from the repository, along with common genealogy metrics, such as clone density, k-volatile, and others. Moreover, one may use OmniCCG in two different ways. The first is via a modern and responsive user interface, in which one can easily track the genealogies in their repository alongside a dashboard of relevant metrics. The second is via a console application that supports local execution. OmniCCG is available as an web application [27] and console application [28].
Denis Sousa, Matheus Paixão, Adriely Silva, Italo Uchoa, Chaiyong Ragkhitwetsagul
MSR5
2026 An Empirical Study of Code Clone Genealogies in Human-AI Collaborative Development
abstract
Code clones consist of two or more identical or similar code snippets. Code clones hurt maintainability by requiring synchronized updates across multiple locations and increasing the risk of inconsistent changes. To understand how clones evolve, the genealogy of code clones captures the evolutionary history of duplicated code snippets by linking them across successive versions of the software system. Since the emergence of Large Language Models (LLMs), software engineering has been reshaped, with code written and evolved differently. This evolution has given rise to coding agents who act as partners to developers. While code clone genealogy is well understood in human-centric development, its evolution in human–agent collaborative projects remains unclear. In this study, we analyze 350 code clone lineages across 6 software projects in which human actively colaborate with coding agents. We observed that humans introduce 85.71% of code clones, whereas agents contribute only 14.29%. Despite similar clone survival rates for both humans (80%) and agents (76%), the maintenance dynamics differ significantly. The analysis of genealogies reveals that humans predominate in maintaining lineages created by agents. These findings highlight that humans remain critical for the evolution of code generated by coding agents.
Denis Sousa, Italo Uchoa, Matheus Paixão, Chaiyong Ragkhitwetsagul, Thiago Lima Matos
MSR2
2026 A Study on Code Clone Lifecycles in Pull Requests Created by AI Agents
abstract
Code clones are fragments of code that are copied and reused within the same or across different codebases, often with minor modifications. Their presence poses significant challenges, as defects or changes in one cloned fragment may require consistent updates across all related clones, negatively affecting software maintainability. Code Clone Lifecycle analysis provides valuable insights into when code clones are introduced and how they evolve during the code review process. Recent advances in Large Language Models (LLMs) have enabled Coding Agents that autonomously create branches, modify code, and submit Pull Requests (PRs). While these agents improve productivity, they also introduce new challenges for managing code clones within PRs. This paper presents an analysis of the Code Clone Lifecycle in agentic PRs hosted on GitHub. Using the NiCad clone detection tool, we analyzed 7,851 PRs created by AI agents from the AiDev dataset. Our results identify 28,425 clones across 497 PRs. Manual validation of a representative sample shows a predominance of Type I (29%) and Type III (46.26%) clones. Among the affected PRs, 93 contain clones restricted to a single commit, 320 exhibit clones recurring across multiple commits, and 84 present both single and recurring occurrences. Overall, the findings indicate that clones tend to persist once introduced, progressing through the PR lifecycle and ultimately being merged into the codebase.
Italo Uchoa, Denis Sousa, Henrique Chuvas, Matheus Paixão, Chaiyong Ragkhitwetsagul, Thiago Lima Matos
MSR1
2024 Code Clone Configuration as a Multi-Objective Search Problem
abstract
Clone detection is an automated process for finding duplicated code within a project’s code base or between online sources. Nowadays, the code cloning community advocates that developers must be aware of the clones they may have in their code bases. In modern clone detection, rank-based tools appear as the ones able to handle the large code corpora that are necessary to identify online clones. However, such tools are sensitive to their parameters, which directly affects their clone detection abilities. Moreover, existing parameter optimization approaches for clone detectors are not meant for rank-based tools. To overcome this issue and facilitate empirical studies of code clones, we introduce Multi-objective Code Clone Configuration, a new approach based on multi-objective optimization to search for an optimal set of parameters for a rank-based clone detection tool. In our empirical evaluation, we ran 3 baseline search algorithms and NSGA-II to assess their performance in this new optimization problem. Additionally, we compared the optimized configurations with the default one. Our results show that NSGA-II was the algorithm that achieved the best performance, finding better configurations than those of the baseline algorithms. Finally, the optimized configurations achieved improvements of 71.08% and 46.29% for our fitness functions.
Denis Sousa, Matheus Paixão, Chaiyong Ragkhitwetsagul, Italo Uchoa
ESEM4