Luciana Lourdes Silva

dblp:93/9384 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-9987-4538ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Unboxing Default Argument Breaking Changes in 1 + 2 data science libraries
João Eduardo Montandon, Luciana Lourdes Silva, Cristiano Politowski, Daniel Prates, Arthur de Brito Bonifácio, Ghizlane El-Boussaidi
J. Syst. Softw.2
2024 Detecting Code Smells using ChatGPT: Initial Insights
abstract
This paper presents initial insights into the effectiveness of ChatGPT in detecting code smells in Java projects. We utilize a large dataset comprising four code smells—Blob, Data Class, Feature Envy, and Long Method—classified into three severity levels. To assess ChatGPT’s proficiency, we employ two different prompts: (i) a generic prompt and (ii) a prompt specifying the smells selected for our research. We evaluate ChatGPT’s abilities using metrics such as precision, recall, and F-measure. Our results reveal that the odds of ChatGPT providing a correct outcome with a specific prompt are 2.54 times higher compared to a generic one. Furthermore, ChatGPT is more effective at detecting smells with critical severity (F-measure = 0.52) than those with minor severity (F-measure = 0.43). Finally, we discuss the implications of our findings and suggest future research directions for leveraging large language models to detect code smells.
Luciana Lourdes Silva, Janio Rosa da Silva, João Eduardo Montandon, Marcus Andrade, Marco Túlio Valente
ESEM1
2023 Unboxing Default Argument Breaking Changes in Scikit Learn
abstract
Machine Learning (ML) has revolutionized the field of computer software development, enabling data-based predictions and decision-making across several domains. Following modern software development practices, developers use third-party libraries—e.g., Scikit Learn, TensorFlow, and PyTorch—to integrate ML-based functionalities into their applications. Due to the complexity inherent in ML techniques, the models available in the APIs of these tools often require an extensive list of arguments to be set up. Library maintainers overcome this issue by defining default values for most of these arguments so developers can use ML models in their client applications effortlessly. By relying on these default arguments, the clients inadvertently depend on the value defined in these parameters to keep running as expected. We interpret this problem as a semantical breaking change variant, which we named Default Argument Breaking Change (DABC). In this work, we leverage 77 DABCs in Scikit Learn—a well-known ML library—and investigate how 194K client applications are vulnerable to them. Our results show that 72 DABCs (93%) are responsible for exposing 67,747 clients (35%). We also detected that most DABCs (61, 79%) involve APIs used in ML model training and model evaluation stages. Finally, we discuss the importance of managing DABCs in third-party ML libraries and provide insights for developers to mitigate the potential impact of these changes in their applications.
João Eduardo Montandon, Luciana Lourdes Silva, Cristiano Politowski, Ghizlane El-Boussaidi, Marco Túlio Valente
SCAM2
2021 What skills do IT companies look for in new developers? A study with Stack Overflow jobs
João Eduardo Montandon, Cristiano Politowski, Luciana Lourdes Silva, Marco Túlio Valente, Fábio Petrillo, Yann-Gaël Guéhéneuc
Inf. Softw. Technol.3
2021 Mining the Technical Roles of GitHub Users
João Eduardo Montandon, Marco Túlio Valente, Luciana Lourdes Silva
Inf. Softw. Technol.3
2020 Is this GitHub project maintained? Measuring the level of maintenance activity of open-source projects
Jailton Coelho, Marco Túlio Valente, Luciano Milen, Luciana Lourdes Silva
Inf. Softw. Technol.4
2019 Identifying experts in software libraries and frameworks among GitHub users
abstract
Software development increasingly depends on libraries and frameworks to increase productivity and reduce time-to-market. Despite this fact, we still lack techniques to assess developers expertise in widely popular libraries and frameworks. In this paper, we evaluate the performance of unsupervised (based on clustering) and supervised machine learning classifiers (Random Forest and SVM) to identify experts in three popular JavaScript libraries: facebook/react, mongodb/node-mongodb, and socketio/socket.io. First, we collect 13 features about developers activity on GitHub projects, including commits on source code files that depend on these libraries. We also build a ground truth including the expertise of 575 developers on the studied libraries, as self-reported by them in a survey. Based on our findings, we document the challenges of using machine learning classifiers to predict expertise in software libraries, using features extracted from GitHub. Then, we propose a method to identify library experts based on clustering feature data from GitHub; by triangulating the results of this method with information available on Linkedin profiles, we show that it is able to recommend dozens of GitHub users with evidences of being experts in the studied JavaScript libraries. We also provide a public dataset with the expertise of 575 developers on the studied libraries.
João Eduardo Montandon, Luciana Lourdes Silva, Marco Túlio Valente
MSR2
2019 Co-change patterns: A large scale empirical study
Luciana Lourdes Silva, Marco Túlio Valente, Marcelo de Almeida Maia
J. Syst. Softw.1
2018 Identifying unmaintained projects in github
abstract
Background: Open source software has an increasing importance in modern software development. However, there is also a growing concern on the sustainability of such projects, which are usually managed by a small number of developers, frequently working as volunteers. Aims: In this paper, we propose an approach to identify GitHub projects that are not actively maintained. Our goal is to alert users about the risks of using these projects and possibly motivate other developers to assume the maintenance of the projects. Method: We train machine learning models to identify unmaintained or sparsely maintained projects, based on a set of features about project activity (commits, forks, issues, etc). We empirically validate the model with the best performance with the principal developers of 129 GitHub projects. Results: The proposed machine learning approach has a precision of 80%, based on the feedback of real open source developers; and a recall of 96%. We also show that our approach can be used to assess the risks of projects becoming unmaintained. Conclusions: The model proposed in this paper can be used by open source users and developers to identify GitHub projects that are not actively maintained anymore.
Jailton Coelho, Marco Túlio Valente, Luciana Lourdes Silva, Emad Shihab
ESEM3
2015 Developers' perception of co-change patterns: An empirical study
abstract
Co-change clusters are groups of classes that frequently change together. They are proposed as an alternative modular view, which can be used to assess the traditional decomposition of systems in packages. To investigate developer's perception of co-change clusters, we report in this paper a study with experts on six systems, implemented in two languages. We mine 102 co-change clusters from the version history of such systems, which are classified in three patterns regarding their projection to the package structure: Encapsulated, Crosscutting, and Octopus. We then collect the perception of expert developers on such clusters, aiming to ask two central questions: (a) what concerns and changes are captured by the extracted clusters? (b) do the extracted clusters reveal design anomalies? We conclude that Encapsulated Clusters are often viewed as healthy designs and that Crosscutting Clusters tend to be associated to design anomalies. Octopus Clusters are normally associated to expected class distributions, which are not easy to implement in an encapsulated way, according to the interviewed developers.
Luciana Lourdes Silva, Marco Túlio Valente, Marcelo de Almeida Maia, Nicolas Anquetil
ICSME1
2014 On the evaluation of an open software engineering course
abstract
Open online courses are a method of online lecturing whose application in education is not bounded by space and location constraints. The successful implementation of open courses requires conceptual changes in how instructors and students behave in open unbounded education environment. There are some emerging open courses for teaching specific topics of Software Engineering. However, it is still limited the knowledge about the best practices for learning Software Engineering processes, methods, and tools in such an open environment. To address this limitation, this paper presents and evaluates an open course for Introduction to Software Engineering. The presented open course has over 250 online students registered and is based on a face-to-face equivalent. The online course is currently composed of 44 video lectures, 160 questions in 16 quizzes, and several discussion topics. We evaluate this course by comparing the students' performance in online vs. face-to-face equivalent courses. Our results indicated that students who had access to online content achieve similar or better performance than students taking only the face-to-face course.
Eduardo Figueiredo 0001, Juliana Alves Pereira, Lucas Garcia, Luciana Lourdes Silva
FIE4