VLDB 2026 Research / reviewers in the wild / expert
Léuson M. P. da Silva
dblp:183/9958 · also Léuson Mario Pedro da Silva
· DBLP profile ↗
11ranked-venue papers
5as first author
8since 2021 · last 2027
0000-0002-9086-9038ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Refactoring with LLMs: Bridging human expertise and machine understanding
Yonnel Chen Kuang Piao, Jean Carlors Paul, Léuson M. P. da Silva, Arghavan Moradi Dakhel, Mohammad Hamdaqa, Foutse Khomh |
Empir. Softw. Eng. | 3 |
| 2025 | A Taxonomy of Inefficiencies in LLM-Generated Python CodeabstractLarge Language Models (LLMs) are widely adopted for automated code generation with promising results. Although prior research has assessed LLM-generated code and identified various quality issues- such as redundancy, poor maintainability, and sub-optimal performance- a systematic understanding and categorization of these inefficiencies remain unexplored. Therefore, we empirically investigate inefficiencies in LLM-generated Python code by state-of-the-art models, i.e., CodeLlama, DeepSeek-Coder, and CodeGemma. To do so, we manually analyze 492 generated Python code snippets in the HumanEval+ dataset. We then construct a taxonomy of inefficiencies in LLM-generated Python code that includes 5 categories (General Logic, Performance, Readability, Maintainability, and Errors) and 19 subcategories of inefficiencies. We validate the obtained taxonomy through an online survey with 58 LLM practitioners and researchers. The surveyed participants affirmed the completeness of the proposed taxonomy, and the relevance and the popularity of the identified code inefficiency patterns. Our qualitative findings indicate that inefficiencies are diverse and interconnected, affecting multiple aspects of code quality, with logic and performance-related inefficiencies being the most frequent and often co-occurring while impacting overall code quality. Our taxonomy provides a structured basis for evaluating the quality of LLM-generated code and guiding future research to improve code generation efficiency. Altaf Allah Abbassi, Léuson M. P. da Silva, Amin Nikanjam, Foutse Khomh |
ICSME | 2 |
| 2025 | Logging requirement for continuous auditing of responsible machine learning-based applications
Patrick Loic Foalem, Léuson M. P. da Silva, Foutse Khomh, Heng Li 0007, Ettore Merlo |
Empir. Softw. Eng. | 2 |
| 2025 | Assessing the adoption of security policies by developers in terraform across different cloud providersabstractCloud computing has become popular thanks to the widespread use of Infrastructure as Code (IaC) tools, allowing the community to manage and configure cloud infrastructure using scripts. However, the scripting process does not automatically prevent practitioners from introducing misconfigurations, vulnerabilities, or privacy risks. As a result, ensuring security relies on practitioners’ understanding and the adoption of explicit policies. To understand how practitioners deal with this problem, we perform an empirical study analyzing the adoption of scripted security best practices present in Terraform files, applied on AWS, Azure, and Google Cloud. We assess the adoption of these practices by analyzing a sample of 812 open-source GitHub projects. We scan each project’s configuration files, looking for policy implementation through static analysis (Checkov and Tfsec). The category Access policy emerges as the most widely adopted in all providers, while Encryption at rest presents the most neglected policies. Regarding the cloud providers, we observe that AWS and Azure present similar behavior regarding attended and neglected policies. Finally, we provide guidelines for cloud practitioners to limit infrastructure vulnerability and discuss further aspects associated with policies that have yet to be extensively embraced within the industry. Alexandre Verdet, Mohammad Hamdaqa, Léuson M. P. da Silva, Foutse Khomh |
Empir. Softw. Eng. | 3 |
| 2025 | Correction to: Assessing the adoption of security policies by developers in terraform across different cloud providers
Alexandre Verdet, Mohammad Hamdaqa, Léuson M. P. da Silva, Foutse Khomh |
Empir. Softw. Eng. | 3 |
| 2025 | LLMs and Stack Overflow discussions: Reliability, impact, and challengesabstractSince its release in November 2022, ChatGPT has shaken up Stack Overflow, the premier platform for developers’ queries on programming and software development. Demonstrating an ability to generate instant, human-like responses to technical questions, ChatGPT has ignited debates within the developer community about the evolving role of human-driven platforms in the age of generative AI. Two months after ChatGPT’s release, Meta released its answer with its own Large Language Model (LLM) called LLaMA: the race was on . We conducted an empirical study analyzing questions from Stack Overflow and using these LLMs to address them. This way, we aim to quantify the reliability of LLMs’ answers and their potential to replace Stack Overflow in the long term; identify and understand why LLMs fail; measure users’ activity evolution with Stack Overflow over time; and compare LLMs together. Our empirical results are unequivocal: ChatGPT and LLaMA challenge human expertise, yet do not outperform it for some domains , while a significant decline in user posting activity has been observed. Furthermore, we also discuss the impact of our findings regarding the usage and development of new LLMs and provide guidelines for future challenges faced by users and researchers. Léuson M. P. da Silva, Jordan Samhi, Foutse Khomh |
J. Syst. Softw. | 1 |
| 2024 | Detecting semantic conflicts with unit testsabstractWhile modern merge techniques, such as 3-way and structured merge, can resolve textual conflicts automatically, they fail when the conflict arises not at the syntactic, but at the semantic level. Detecting such semantic conflicts requires understanding the behavior of the software, which is beyond the capabilities of most existing merge tools. Although semantic merge tools have been proposed, they are usually based on heavyweight static analyses, or need explicit specifications of program behavior. In this work, we take a different route and propose SAM (SemAntic Merge), a semantic merge tool based on the automated generation of unit tests that are used as partial specifications of the changes to be merged, and that drive the detection of unwanted behavior changes (conflicts) when merging software. To evaluate SAM’s feasibility for detecting conflicts, we perform an empirical study relying on a dataset of more than 80 pairs of changes integrated to common class elements (constructors, methods, and fields) from 51 merge scenarios. We also assess how the four unit test generation tools used by SAM individually contribute to conflict identification. Our results show that SAM performs best when combining only the tests generated by Differential EvoSuite and EvoSuite, and using our proposed testability transformations (nine detected conflicts out of 29). These results reinforce previous findings about the potential of using test-case generation to detect conflicts as a method that is versatile and requires only limited deployment effort in practice. Léuson M. P. da Silva, Paulo Borba, Toni Maciel, Wardah Mahmood, Thorsten Berger, João Moisakis, Aldiberg Gomes, Vinícius Leite |
J. Syst. Softw. | 1 |
| 2022 | Build conflicts in the wildabstractAbstract When collaborating, developers often create and change software artifacts without being fully aware of team members' work. While such independence is essential for increasing development productivity, it might also result in conflicts when integrating developers' contributions. To better understand the conflicts revealed by failures when building integrated code, we investigate their frequency and structure and adopted resolution patterns in 451 open‐source Java projects. To detect such build conflicts, we select merge scenarios from git repositories, parse their Travis build logs, and check whether the build error messages are related to the merged changes. We find and classify 239 build conflicts and their resolution patterns. Most build conflicts are caused by missing declarations removed or renamed by one developer but referenced by another developer. Conflicts caused by renaming are often resolved by updating the missing reference, whereas removed declarations are often reintroduced. Most fix commits are authored by one of the merge scenario contributors. Based on our catalogue of conflict causes, awareness tools could alert developers about the risk of conflict situations. Program repair tools could benefit from our catalogue of resolution patterns to automatically fix conflicts; we illustrate that with a proof of concept implementation of a tool that fixes conflicts. Léuson M. P. da Silva, Paulo Borba, Arthur Pires |
J. Softw. Evol. Process. | 1 |
| 2020 | Detecting Semantic Conflicts via Automated Behavior Change DetectionabstractBranching and merging are common practices in collaborative software development. They increase developer productivity by fostering teamwork, allowing developers to independently contribute to a software project. Despite such benefits, branching and merging comes at a cost-the need to merge software and to resolve merge conflicts, which often occur in practice. While modern merge techniques, such as 3-way or structured merge, can resolve many such conflicts automatically, they fail when the conflict arises not at the syntactic, but the semantic level. Detecting such conflicts requires understanding the behavior of the software, which is beyond the capabilities of most existing merge tools. As such, semantic conflicts can only be identified and fixed with significant effort and knowledge of the changes to be merged. While semantic merge tools have been proposed, they are usually heavyweight, based on static analysis, and need explicit specifications of program behavior. In this work, we take a different route and explore the automated creation of unit tests as partial specifications to detect unwanted behavior changes (conflicts) when merging software.We systematically explore the detection of semantic conflicts through unit-test generation. Relying on a ground-truth dataset of 38 software merge scenarios, which we extracted from GitHub, we manually analyzed them and investigated whether semantic conflicts exist. Next, we apply test-generation tools to study their detection rates. We propose improvements (code transformations) and study their effectiveness, as well as we qualitatively analyze the detection results and propose future improvements. For example, we analyze the generated test suites for false-negative cases to understand why the conflict was not detected. Our results evidence the feasibility of using test-case generation to detect semantic conflicts as a method that is versatile and requires only limited deployment effort in practice, as well as it does not require explicit behavior specifications. Léuson M. P. da Silva, Paulo Borba, Wardah Mahmood, Thorsten Berger, João Moisakis |
ICSME | 1 |
| 2018 | Analyzing conflict predictors in open-source Java projectsabstractIn collaborative development environments integration conflicts occur frequently. To alleviate this problem, different awareness tools have been proposed to alert developers about potential conflicts before they become too complex. However, there is not much empirical evidence supporting the strategies used by these tools. Learning about what types of changes most likely lead to conflicts might help to derive more appropriate requirements for early conflict detection, and suggest improvements to existing conflict detection tools. To bring such evidence, in this paper we analyze the effectiveness of two types of code changes as conflict predictors. Namely, editions to the same method, and editions to directly dependent methods. We conduct an empirical study analyzing part of the development history of 45 Java projects from GitHub and Travis CI, including 5,647 merge scenarios, to compute the precision and recall for the conflict predictors aforementioned. Our results indicate that the predictors combined have a precision of 57.99% and a recall of 82.67%. Moreover, we conduct a manual analysis which provides insights about strategies that could further increase the precision and the recall. Paola R. G. Accioly, Paulo Borba, Léuson M. P. da Silva, Guilherme Cavalcanti |
MSR | 3 |
| 2017 | Autonomy in Software Engineering: A Preliminary Study on the Influence of Education Level and Professional ExperienceabstractContext: Software development process is executed by professionals with different roles, who are responsible for distinct activities. These roles can have different degrees of autonomy depending on some factors, such as the adopted process and hierarchy. Goal: This study aims to identify what factors can impact autonomy and also investigate how autonomy is given to an employee based on two main factors: education level and professional experience. Methodology: Initially, a survey was carried out to understand how autonomy is perceived by 102 software engineers, as well as by 83 professionals from other areas. The next step was applying semi-structured interviews with software engineers to find a better understanding of the quantitative findings. Results: In general, education level and professional experience do not have an impact on autonomy. Only when autonomy is evaluated from the education level perspective, there is a significant difference among the respondents. During the interviews, we also could identify some topics that respondents mentioned which were related to autonomy. For example, the experience that software engineer has in a current project and the development process adopted by the company influence how autonomy is perceived. Conclusion: While professional qualification and experience are not directly related to autonomy, the lack of process and the amount of work experience on specific projects seem to be relevant factors to be aware of. Léuson M. P. da Silva, Alberto Trindade Tavares, Victor Afonso dos Santos Ferreira, Alex Juvencio Costa, Gabriel Ibson de Souza, Claudio Jose Antunes Salgueiro Magalhaes, Fabio Q. B. da Silva |
ESEM | 1 |