Sayma Sultana

dblp:264/7972 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-8316-2560ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Automated Identification of Sexual Orientation and Gender Identity Discriminatory Texts from Issue Comments
abstract
In an industry dominated by straight men, many developers representing other gender identities and sexual orientations often encounter hateful or discriminatory messages. Such communications pose barriers to participation for women and LGBTQ+ persons. Due to sheer volume, manual inspection of all communications for discriminatory communication is infeasible for a large-scale Free Open Source Software (FLOSS) community. To address this challenge, this study proposes an automated mechanism to identify Sexual Orientation and Gender Identity Discriminatory (SGID) texts in software developers’ communications. On this goal, we trained and evaluated SGID4SE (Sexual orientation and Gender Identity Discriminatory text identification for (4) Software Engineering texts), a supervised learning-based tool. SGID4SE incorporates six preprocessing steps and 10 state-of-the-art algorithms. SGID4SE employs six distinct strategies to enhance the performance of the minority class. We empirically evaluated each strategy and identified an optimum configuration for each algorithm. In our ten-fold cross-validation-based evaluations, a BERT-based model achieves the best performance with 85.9% precision, 80.0% recall, and 82.9% F1-score for the SGID class. This model achieves 95.7% accuracy and a Matthews Correlation Coefficient of 80.4%. Our dataset and tool establish a foundation for further research in this direction.
Sayma Sultana, Jaydeb Sarker, Farzana Israt, Rajshakhar Paul, Amiangshu Bosu
ACM Trans. Softw. Eng. Methodol.1
2025 Beyond Binary Moderation: Identifying Fine-Grained Sexist and Misogynistic Behavior on GitHub with Large Language Models
abstract
Background: Sexist and misogynistic behavior significantly hinders inclusion in technical communities like GitHub, causing developers, especially minorities, to leave due to subtle biases and microaggressions. Current moderation tools primarily rely on keyword filtering or binary classifiers, limiting their ability to detect nuanced harm effectively. Aims: This study introduces a fine-grained, multi-class classification framework that leverages instruction-tuned Large Language Models (LLMs) to identify twelve distinct categories of sexist and misogynistic comments on GitHub. Method: We utilized an instruction-tuned LLM-based framework with systematic prompt refinement across 20 iterations, evaluated on 1,440 labeled GitHub comments across twelve sexism/misogyny categories. Model performances were rigorously compared using precision, recall, F1-score, and the Matthews Correlation Coefficient (MCC). Results: Our optimized approach (GPT-4o with Prompt 19) achieved an MCC of 0.501, significantly outperforming baseline approaches. While this model had low false positives, it struggled to interpret nuanced, contextdependent sexism and misogyny reliably. Conclusion: Welldesigned prompts with clear definitions and structured outputs significantly improve the accuracy and interpretability of sexism detection, enabling precise and practical moderation on developer platforms like GitHub.
Tanni Dev, Sayma Sultana, Amiangshu Bosu
ESEM2
2025 From First Patch to Long-Term Contributor: Evaluating Onboarding Recommendations for OSS Newcomers
abstract
Attracting and retaining a steady stream of new contributors is crucial to ensuring the long-term survival of open-source software (OSS) projects. However, there are two key research gaps regarding recommendations for onboarding new contributors to OSS projects. First, most of the existing recommendations are based on a limited number of projects, which raises concerns about their generalizability. If a recommendation yields conflicting results in a different context, it could hinder a newcomer's onboarding process rather than help them. Second, it's unclear whether these recommendations also apply to experienced contributors. If certain recommendations are specific to newcomers, continuing to follow them after their initial contributions are accepted could hinder their chances of becoming long-term contributors. To address these gaps, we conducted a two-stage mixed-method study. In the first stage, we conducted a Systematic Literature Review (SLR) and identified 15 task-related actionable recommendations that newcomers to OSS projects can follow to improve their odds of successful onboarding. In the second stage, we conduct a large-scale empirical study of five Gerrit-based projects and 1,155 OSS projects from GitHub to assess whether those recommendations assist newcomers’ successful onboarding. Our results suggest that four recommendations positively correlate with newcomers’ first patch acceptance in most contexts. Four recommendations are context-dependent, and four indicate significant negative associations for most projects. Our results also found three newcomer-specific recommendations, which OSS joiners should abandon at non-newcomer status to increase their odds of becoming long-term contributors.
Asif Kamal Turzo, Sayma Sultana, Amiangshu Bosu
IEEE Trans. Software Eng.2
2023 ToxiSpanSE: An Explainable Toxicity Detection in Code Review Comments
abstract
Background: The existence of toxic conversations in open-source platforms can degrade relationships among software developers and may negatively impact software product quality. To help mitigate this, some initial work has been done to detect toxic comments in the Software Engineering (SE) domain. Aims: Since automatically classifying an entire text as toxic or non-toxic does not help human moderators to understand the specific reason(s) for toxicity, we worked to develop an explainable toxicity detector for the SE domain. Method: Our explainable toxicity detector can detect specific spans of toxic content from SE texts, which can help human moderators by automatically highlighting those spans. This toxic span detection model, ToxiSpanSE, is trained with the 19,651 code review (CR) comments with labeled toxic spans. Our annotators labeled the toxic spans within 3,757 toxic CR samples. We explored several types of models, including one lexicon-based approach and five different transformer-based encoders. Results: After an extensive evaluation of all models, we found that our fine-tuned RoBERTa model achieved the best score with 0.88$F1$, 0.87 precision, and 0.93 recall for toxic class tokens, providing an explainable toxicity classifier for the SE domain. Conclusion: Since ToxiSpanSE is the first tool to detect toxic spans in the SE domain, this tool will pave a path to combat toxicity in the SE community.
Jaydeb Sarker, Sayma Sultana, Steven R. Wilson 0001, Amiangshu Bosu
ESEM2
2023 Code reviews in open source projects : how do gender biases affect participation and outcomes?
Sayma Sultana, Asif Kamal Turzo, Amiangshu Bosu
Empir. Softw. Eng.1
2022 Identification and Mitigation of Gender Biases to Promote Diversity and Inclusion among Open Source Communities
abstract
Contemporary software development organizations are dominated by straight males and lack diversity. As a result, people from other demographic such as women and LGBTQ+ often encounter bias, sexism, and misogyny. Due to negative experiences, many women switch careers. Therefore, biases pose barriers to promote diversity and inclusion. To get benefits from diverse pools of talents and reduce the attrition rate of minorities, we need to identify the degree and effect of various biases and develop mitigation strategies. Therefore, my dissertation study aims at promoting diversity and inclusion among software development organizations by identifying the manifestation, magnitude, and frequency of various gender biases. For this purpose, I plan to investigate i) the effect of gender of the contributors in the code review process of Free/Libre Open Source Software (FLOSS) projects, ii) the frequency of different dimensions of gender bias and their effect, and iii) develop a tool to identify sexist and misogynistic and derogatory (SMD) texts.
Sayma Sultana
ASE1
2022 Identifying Sexism and Misogyny in Pull Request Comments
abstract
Being extremely dominated by men, software development organizations lack diversity. People from other groups often encounter sexist, misogynistic, and discriminatory (SMD) speech during communication. To identify SMD contents, I aim to build an automatic misogyny identification (AMI) tool for the domain of software developers. On this goal, I built a dataset of 10,138 pull request comments mined from Github based on a keyword-based selection, followed by manual validation. Using ten-fold cross-validation, I evaluated ten machine learning algorithms for automatic identification. The best performing model achieved 80% precision, 67.07% recall, 72.5% f-score, and 95.96% accuracy.
Sayma Sultana
ASE1
2021 A Rubric to Identify Misogynistic and Sexist Texts from Software Developer Communications
abstract
Background: As contemporary software development organizations are dominated by males, occurrences of misogynistic and sexist remarks are abundant in many communities. Such remarks are barriers to promoting diversity and inclusion in the software engineering (SE) domain.
Sayma Sultana, Jaydeb Sarker, Amiangshu Bosu
ESEM1