Liviu Berciu

dblp:363/1385 · also Liviu-Marian Berciu · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0009-7501-2851ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A long-term exploratory study of source code quality issues in open-source Python projects
abstract
Abstract Empirical research targeting software quality resulted in a consistent body of work, especially with the help of automated tools that allow researchers to mine large amounts of data. However, we find that many of these efforts provide cross-sectional or only short-term longitudinal analyses. Furthermore, they are often focused on a single and in most cases statically typed language such as Java. In the present paper, we aim to broaden the horizon of existing efforts by exploring the composition, distribution, and evolution of source code quality issues in complex, open-source Python projects. We employ the SonarQube static analysis tool on a dataset comprised of 3656 individual releases of 57 Python projects. We explore the impact of these issues on software maintainability, reliability, and security. We investigate the evolution of these issues over the long term and compare our results with those in the literature. We compare our findings with existing research targeting both Python and Java; for the latter, we investigate the impact the development language has on the type and distribution of detected issues. Our study data are published and open source to help replicate our investigation and contribute to building open and large-scale data sets for research.
Liviu Berciu, Simona Motogna, Arthur-Jozsef Molnar
Softw. Qual. J.1
2025 Exploring the relation between source code commit information and SonarQube issues
abstract
Commit classification has emerged as a practical approach to improve software quality, providing a systematic way to interpret development activities, measure their impact, and enhance overall software quality. The Conventional Commit Specification offers a structured, detailed framework to classify commits beyond traditional maintenance categories. In this study, we explore how conventional commit specification categories are distributed between software releases. We analyze 2,600 software release pairs of 57 Python open-source projects and employ large language models to label 90,318 commits. We identify which commit types are predominantly responsible for introducing SonarQube issues, and explore the relationship between commit types and clean-code attributes affected by quality issues. We publish our analysis dataset to enable replicating or extending our study. We find that most of the commit types are related to documentation and bug fixing, and the existence of a trend of focusing on improving code quality by tackling common code smells and adhering to best practices. We discuss the threats to the validity of our study and identify avenues for further exploration.
Liviu Berciu, Simona Motogna, Oskar Picus, Arthur-Jozsef Molnar
KES1
2024 The Python Software Quality Dataset
abstract
With Python's ascension as a dominant program-ming language, particularly in the fields of artificial intelligence and data science, the need for comprehensive datasets focusing on software quality within Python projects has become increasingly noticeable. This study introduces a detailed dataset designed to address this gap, enriching academic resources in software engineering. The dataset encompasses a wide array of software quality metrics on up to 80 projects, including 51.765.853 Sonar-Qube issues, 268.506 SonarQube code quality metrics, 11.915 software refactoring records, and 155.127 pairs of bug-inducing and bug-fixing commits, along with 863.931 GitHub issue tracker entries. This extensive collection serves as a versatile tool for various research activities, enabling analysis of the relationships between technical debt and software refactorings, correlations be-tween refactoring processes and bug resolution, and their overall impact on software maintainability and reliability. By offering a comprehensive and multifaceted dataset, this study significantly contributes to understanding and improving software quality in Python projects.
Vasilica-Andreea Moldovan, Liviu Berciu, Rares-Danut Patcas
SEAA2
2024 Artificial Intelligence Methods in Software Refactoring: A Systematic Literature Review
abstract
Refactoring is an important process in software engineering, aiming to improve code quality without altering the behavior. This article presents a systematic review of the literature (SLR) on artificial intelligence in the domain of software refactoring. Following a rigorous methodology consisting of data extraction, snowballing techniques, and manual validation, we created a dataset consisting of 156 articles. The focus of the investigation was to identify the refactoring stages that are addressed. The results show that, as research type, most of the contributions propose solutions, while other forms of research such as evaluation, validation and experience are less represented in publications. Refactoring detection represents the highest interest in research contributions, while other refactoring stages, such as prioritization or testing are less investigated. The most commonly used AI methods include Random Forests, Genetic Algorithms, SVM, CNNs and Decision Trees. Based on this literature review, we have identified research trends and opportunities for future research.
Simona Motogna, Liviu Berciu, Vasilica-Andreea Moldovan
SEAA2