Fabio Santos

dblp:289/0005 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-8069-3158ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 The Shifting Sands of Toxicity: The Evolving Nature of Interpersonal Challenges in Open Source
abstract
Background: The sustainability of Open Source Software (OSS) projects relies on attracting and retaining contributors. Interpersonal challenges, whether experienced or witnessed, can discourage participation, alter behavior, or drive contributors away. Aims: This study examines how interpersonal challenges in OSS communities persist over time and explores their behavioral consequences on contributors' decisions and actions. Method: We analyze data from two large GitHub Open Source surveys conducted in$2017(\mathrm{n}=5,495)$and$2024(\mathrm{n}=8,452)$, evaluating changes in reported interpersonal challenges (RQ1) and differential consequences of exposure between 2017 and 2024 (RQ2). Results: Our findings reveal a significant increase in reported interpersonal challenges in 2024 compared to 2017. Contributors more frequently reported severe challenges such as threats of violence, impersonation, sustained harassment, stalking, and doxxing. The behavioral impact has shifted: experiencing rudeness, stalking, and name-calling became strongly linked to stopping contributions, adopting pseudonyms, working privately, and avoiding offline events. Witnessing harmful behaviors like name-calling and impersonation also became stronger predictors of working privately or advocating for Codes of Conduct. These trends show toxicity is not only more pervasive but increasingly damaging to OSS participation and community health. Conclusions: Results highlight a concerning rise in interpersonal challenges within OSS communities, with rudeness emerging as the most impactful. The growing influence of toxic behaviors on contributors' decisions to withdraw, conceal identities, isolate collaboration, and avoid offline engagement underscores the urgent need for stronger, proactive community support. Sustaining healthy OSS projects requires both technical excellence and deliberate investment in social infrastructure to foster respectful collaboration spaces.
Sarthak Bharadwaj, Fabio Santos, Bianca Trinkenreich
ESEM2
2025 Contribution History as a Key Feature in OSS Task Recommendation: An LLM-Based Empirical Study
abstract
Background: Open-source software (OSS) projects often struggle to efficiently assign issues to contributors whose skills align with task requirements. Without targeted recommendations, contributors may overlook suitable issues, leading to delayed resolutions and reduced engagement. Seeking to mitigate this barrier, previous studies proposed tagging of issues with categories of libraries as a proxy for the skills sufficient to solve them in an attempt to support managers in performing the allocation. Notwithstanding the advances, if contributors are overconfident, they still might pick an issue to solve beyond their abilities. Aims: In this paper, we present a fully automated, AI-powered issue recommendation system that integrates past contribution history with skill-based matching. We mine public GitHub repositories to extract contributor skills using commit histories and issue resolutions, and infer issue requirements using both traditional techniques and Large Language Models (LLMs). Method: We evaluate three matching strategies-TFIDF, sentence-BERT (s-BERT), and LLM-based approaches. We also explore the use of a canonical skill superset for standardizing skill representations. Results: We find that the simple TF-IDF model outperforms more complex methods, achieving a top15 accuracy of 70% and that historical contribution data is a significant feature for OSS issue assignment. Conclusion: Lightweight lexical methods remain highly effective in specific tasks. Therefore, integrating it with other features might improve performance. This work contributes a scalable framework for personalized issue recommendation that supports diverse OSS environments and enhances contributor-task alignment.
Md Abdul Hannan, Mohammad Habibullah Rakib, Khondaker Masfiq Reza, Fabio Santos
ESEM4
2025 Beyond the Job Posting: What Hiring Managers Seek in Entry-Level Software Engineering Candidates
abstract
Background: With limited job openings and growing selectivity in tech, understanding the skills hiring managers expect from entry-level Software Engineering candidates is essential. While job postings emphasize technical abilities, success often hinges on non-technical skills, as interpersonal and adaptive abilities. Yet, mismatches between candidate profiles and employer expectations lead to inefficient hiring and missed opportunities. Aims: We investigated which skills hiring managers expect from early-career Software Engineering (SE) candidates and how those skills are described in job postings. Method: To achieve this goal, we conducted an exploratory case study in Hewlett Packard Enterprise (HPE), a global technology company. We interviewed 12 managers, collected seven HPE job postings related to entry-level positions in Software Engineering-related roles, and employed mixed methods to analyze the data. Results: Our findings show that hiring managers place strong emphasis on non-technical expectations when evaluating entry-level candidates. They commonly believe that technical competencies can be taught on the job, provided the candidate demonstrates the right nontechnical attributes. The non-technical expectations included adaptability, proactivity, dependability, self-direction, passion, collaboration, communication, problem solving, courage, selfawareness, self-confidence, willingness to learn, and having a process-oriented mindset. Conclusions: To secure their first job, Software Engineering students should invest not only in technical competencies but also in developing non-technical skills. Our findings show that hiring managers prioritize these skills, even though they are often underrepresented in job postings. Making such expectations more visible could help bridge the gap between candidates and employers, leading to more efficient and effective hiring.
Spencer Baloga Loufek, Fabio Santos, Bianca Trinkenreich
ESEM2
2025 Is LLM-Generated Code More Maintainable & Reliable Than Human-Written Code?
abstract
Background: The rise of Large Language Models (LLMs) in software development has opened new possibilities for code generation. Despite the widespread use of this technology, it remains unclear how well LLMs generate code solutions in terms of software quality and how they compare to humanwritten code. Aims: This study compares the internal quality attributes of LLM-generated and human-written code. Method: Our empirical study integrates datasets of coding tasks, three LLM configurations (zero-shot, few-shot, and fine-tuning), and SonarQube to assess software quality. The dataset comprises Python code solutions across three difficulty levels: introductory, interview, and competition. We analyzed key code quality metrics, including maintainability and reliability, and the estimated effort required to resolve code issues. Results: Our analysis shows that LLM-generated code has fewer bugs and requires less effort to fix them overall. Interestingly, fine-tuned models reduced the prevalence of high-severity issues, such as blocker and critical bugs, and shifted them to lower-severity categories, but decreased the model's performance. In competition-level problems, the LLM solutions sometimes introduce structural issues that are not present in human-written code. Conclusion: Our findings provide valuable insights into the quality of LLM-generated code; however, the introduction of critical issues in more complex scenarios highlights the need for a systematic evaluation and validation of LLM solutions. Our work deepens the understanding of the strengths and limitations of LLMs for code generation.
Alfred Santa Molison, Marcia Moraes, Glaucia Melo dos Santos, Fabio Santos, Wesley K. G. Assunção
ESEM4
2025 Applying large language models to issue classification: Revisiting with extended data and new models
Gabriel Aracena, Kyle Luster, Fabio Santos, Igor Steinmacher, Marco Aurélio Gerosa
Sci. Comput. Program.3
2024 Predicting Attrition among Software Professionals: Antecedents and Consequences of Burnout and Engagement
abstract
In this study of burnout and engagement, we address three major themes. First, we offer a review of prior studies of burnout among IT professionals and link these studies to the Job Demands-Resources (JD-R) model. Informed by the JD-R model, we identify three factors that are organizational job resources and posit that these (a) increase engagement and (b) decrease burnout. Second, we extend the JD-R by considering software professionals’ intention to stay as a consequence of these two affective states, burnout and engagement. Third, we focus on the importance of factors for intention to stay, and actual retention behavior. We use a unique dataset of over 13,000 respondents at one global IT organization, enriched with employment status 90 days after the initial survey. Leveraging partial-least squares structural quation modeling and machine learning, we find that the data mostly support our theoretical model, with some variation across different subgroups of respondents. An importance-performance map analysis suggests that managers may wish to focus on interventions regarding burnout as a predictor of intention to leave. The Machine Learning model suggests that engagement and opportunities to learn are the top two most important factors that explain whether software professionals leave an organization.
Bianca Trinkenreich, Fabio Santos, Klaas-Jan Stol
ACM Trans. Softw. Eng. Methodol.2
2023 Tell Me Who Are You Talking to and I Will Tell You What Issues Need Your Skills
abstract
Selecting an appropriate task is challenging for newcomers to Open Source Software (OSS) projects. To facilitate task selection, researchers and OSS projects have leveraged machine learning techniques, historical information, and textual analysis to label tasks (a.k.a. issues) with information such as the issue type and domain. These approaches are still far from mainstream adoption, possibly because of a lack of good predictors. Inspired by previous research, we advocate that label prediction might benefit from leveraging metrics derived from communication data and social network analysis (SNA) for issues in which social interaction occurs. Thus, we study how these "social metrics" can improve the automatic labeling of open issues with API domains—categories of APIs used in the source code that solves the issue—which the literature shows that newcomers to the project consider relevant for task selection. We mined data from OSS projects’ repositories and organized it in periods to reflect the seasonality of the contributors’ project participation. We replicated metrics from previous work and added social metrics to the corpus to predict API-domain labels. Social metrics improved the performance of the classifiers compared to using only the issue description text in terms of precision, recall, and F-measure. Precision (0.922) increased by 15.82% and F-measure (0.942) by 15.89% for a project with high social activity. These results indicate that social metrics can help capture the patterns of social interactions in a software project and improve the labeling of issues in an issue tracker.
Fabio Santos, Jacob Penney, João Felipe Pimentel, Igor Scaliante Wiese, Igor Steinmacher, Marco Aurélio Gerosa
MSR1
2023 GiveMeLabeledIssues: An Open Source Issue Recommendation System
abstract
Developers often struggle to navigate an Open Source Software (OSS) project’s issue-tracking system and find a suitable task. Proper issue labeling can aid task selection, but current tools are limited to classifying the issues according to their type (e.g., bug, question, good first issue, feature, etc.). In contrast, this paper presents a tool (GiveMeLabeledIssues) that mines project repositories and labels issues based on the skills required to solve them. We leverage the domain of the APIs involved in the solution (e.g., User Interface (UI), Test, Databases (DB), etc.) as a proxy for the required skills. GiveMeLabeledIssues facilitates matching developers’ skills to tasks, reducing the burden on project maintainers. The tool obtained a precision of 83.9% when predicting the API domains involved in the issues. The replication package contains instructions on executing the tool and including new projects. A demo video is available at https://www.youtube.com/watch?v=ic2quUue7i8
Joseph Vargovich, Fabio Santos, Jacob Penney, Marco Aurélio Gerosa, Igor Steinmacher
MSR2
2023 Tag that issue: applying API-domain labels in issue tracking systems
Fabio Santos, Joseph Vargovich, Bianca Trinkenreich, Ítalo Santos, Jacob Penney, Ricardo Britto 0001, João Felipe Pimentel, Igor Scaliante Wiese, Igor Steinmacher, Anita Sarma, Marco Aurélio Gerosa
Empir. Softw. Eng.1
2022 How to Choose a Task? Mismatches in Perspectives of Newcomers and Existing Contributors
abstract
[Background] Selecting an appropriate task is challenging for Open Source Software (OSS) project newcomers and a variety of strategies can help them in this process. [Aims] In this research, we compare the perspective of maintainers, newcomers, and existing contributors about the importance of strategies to support this process. Our goal is to identify possible gulfs of expectations between newcomers who are meant to be helped and contributors who have to put effort into these strategies, which can create friction and impede the usefulness of the strategies. [Method] We interviewed maintainers (n=17) and applied inductive qualitative analysis to derive a model of strategies meant to be adopted by newcomers and communities. Next, we sent a questionnaire (n=64) to maintainers, frequent contributors, and newcomers, asking them to rank these strategies based on their importance. We used the Schulze method to compare the different rankings from the different types of contributors. [Results] Maintainers and contributors diverged in their opinions about the relative importance of various strategies. The results suggest that newcomers want a better contribution process and more support to onboard, while maintainers expect to solve questions using the available communication channels. [Conclusions] The gaps in perspectives between newcomers and existing contributors create a gulf of expectation. OSS communities can leverage our results to prioritize the strategies considered the most important by newcomers.
Fabio Santos, Bianca Trinkenreich, João Felipe Pimentel, Igor Scaliante Wiese, Igor Steinmacher, Anita Sarma, Marco Aurélio Gerosa
ESEM1
2021 Can I Solve It? Identifying APIs Required to Complete OSS Tasks
abstract
Open Source Software projects add labels to open issues to help contributors choose tasks. However, manually labeling issues is time-consuming and error-prone. Current automatic approaches for creating labels are mostly limited to classifying issues as a bug/non-bug. In this paper, we investigate the feasibility and relevance of labeling issues with the domain of the APIs required to complete the tasks. We leverage the issues' description and the project history to build prediction models, which resulted in precision up to 82% and recall up to 97.8%. We also ran a user study (n=74) to assess these labels' relevancy to potential contributors. The results show that the labels were useful to participants in choosing tasks, and the API-domain labels were selected more often than the existing architecture-based labels. Our results can inspire the creation of tools to automatically label issues, helping developers to find tasks that better match their skills.
Fabio Santos, Igor Scaliante Wiese, Bianca Trinkenreich, Igor Steinmacher, Anita Sarma, Marco Aurélio Gerosa
MSR1