VLDB 2026 Research / reviewers in the wild / expert
Md. Rakibul Islam 0002
dblp:136/5279-2
· DBLP profile ↗
14ranked-venue papers
10as first author
7since 2021 · last 2026
0009-0009-9104-874XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 10 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DBSecQA: A Curated Dataset of Developer Discussions on Database Security from Stack ExchangeabstractDatabase security has become increasingly critical in modern software development, yet developers continue to face significant challenges in implementing and maintaining secure database environments. To support research in this domain, we present DBSecQA, a comprehensive, curated dataset of 22,568 database security-related questions and discussions collected from Stack Overflow and Information Security Stack Exchange. The dataset includes complete metadata (tags, scores, view counts, timestamps, accepted answers) and sophisticated topic categorization using Latent Dirichlet Allocation (LDA) optimized with genetic algorithms, resulting in 20 distinct security topics grouped into 5 main categories. Our dataset enables multiple research directions, including empirical software engineering studies, machine learning applications, educational research, and security tool development. To demonstrate the dataset’s utility, we present preliminary analyses revealing that encryption and authentication represent 41% of all discussions, with encryption-related topics showing the highest percentage of unanswered questions (62.93%). The dataset, along with all collection and processing scripts, is publicly available under an open license, following FAIR data principles for long-term accessibility and reusability. Md. Rakibul Islam 0002, Farha Kamal, Md Humaun Kabir, Md Murad Sharif |
MSR | 1 |
| 2025 | An Empirical Study of Database Security Topics on Technical Social Forums of Software DevelopersabstractDatabase security has become increasingly critical in modern software development, yet developers continue to face significant challenges in implementing and maintaining secure database environments. This paper presents a comprehensive empirical analysis of 22,568 database security-related discussions from Stack Overflow and Information Security Stack Exchange, offering insights into the practical challenges faced by developers in the field. Through sophisticated topic modeling using Latent Dirichlet Allocation (LDA) optimized with genetic algorithms, we identify and analyze the primary topics of discussion, types of questions asked, and areas that prove most challenging based on response patterns. Our findings reveal that encryption and authentication remain the most complex aspects, accounting for 41% of all discussions, with encryption-related topics showing the highest percentage of unanswered questions (62.93%). Analysis of question types demonstrates that 70% of queries focus on implementation guidance (“how-to” questions), highlighting a significant need for practical solutions. Notably, we found no correlation between a topic’s popularity and its difficulty level, suggesting that some critical security challenges may be underserved by the developer community. These insights have important implications for practitioners, researchers, and educators in the database security field, pointing to areas where additional resources and focus are needed. Our study provides valuable direction for improving security practices, enhancing documentation, and developing more effective educational resources for database security implementation. Md. Rakibul Islam 0002, Youngeun Jo |
EASE | 1 |
| 2025 | Robust or Overfitted? Investigating the Generalization of Pretrained Models in Requirement ClassificationabstractBackground: Accurate classification of nonfunctional requirements (NFRs) is essential for aligning stakeholder expectations with system design and ensuring software quality. While transformer-based models such as PRCBERT and NORBERT have achieved high performance in supervised settings, their generalizability across diverse sources of requirements remains largely unexplored. In practice, requirements originate from heterogeneous platforms, ranging from structured specification documents to informal developer discussions on forums like Stack Overflow. Aim: This study provides the first comprehensive, bidirectional cross-dataset evaluation of domain-specific, embedding-based, and promptbased large language models (LLMs) for NFR classification across two contrasting platforms: PROMISE (structured) and NFR-SO (informal). Method: We evaluate domain-specific finetuned models, sentence embedding models, and prompt-based LLMs (including GPT-4o) in both zero-shot and few-shot settings. Performance is measured both in-domain and in crossplatform transfer scenarios to assess generalization with minimal or no labeled data. Results: Domain-specific fine-tuned models, although effective in-domain, exhibit substantial performance degradation when transferred across platforms. In contrast, LLMs-particularly GPT-4o in few-shot mode-consistently outperform other approaches in cross-platform scenarios, achieving strong generalization with minimal labeled data. In zero-shot mode, GPT-4o also demonstrates robust performance without any supervision. Conclusions: Traditional supervised models face limitations in cross-platform NFR classification. Prompt-based LLMs offer a scalable, low-supervision solution for diverse requirement sources. Farha Kamal, Md. Rakibul Islam 0002 |
ESEM | 2 |
| 2024 | A Four-Dimension Gold Standard Dataset for Opinion Mining in Software EngineeringabstractWe present the first four-dimension gold standard dataset to advance opinion mining focused on the software engineering domain. Through a well-defined sampling and annotation strategy leveraging multiple coders, we construct a corpus of 2,000 Stack Overflow posts labeled with four dimensions/tuples, including sentiments, polar facts, aspects, and named entities. This multidimensional ground truth dataset opens up new research opportunities for opinion mining in domain-adapted NLP tools for software engineering by capturing existing relationships between extracted elements at a more granular level. It also facilitates investigating the effects of sentiments in the developers' social forums. Md. Rakibul Islam 0002, Md. Fazle Rabbi, Youngeun Jo, Arifa I. Champa, Ethan Young, Camden Wilson, Gavin Scott, Minhaz Fahim Zibran |
MSR | 1 |
| 2024 | AI Writes, We Analyze: The ChatGPT Python Code SagaabstractIn this study, we quantitatively analyze 1,756 AI-written Python code snippets in the DevGPT dataset and evaluate them for quality and security issues. We systematically distinguish the code snippets as either generated by ChatGPT from scratch (ChatGPT-generated) or modified user-provided code (ChatGPT-modified). The results reveal that ChatGPT-modified code more frequently displays quality issues compared to ChatGPT-generated code. The findings provide insights into the inherent limitations of AI-written code and emphasize the need for scrutiny before integrating such pieces of code into software systems. Md. Fazle Rabbi, Arifa I. Champa, Minhaz Fahim Zibran, Md. Rakibul Islam 0002 |
MSR | 4 |
| 2024 | On the Taxonomy of Developers' Discussion Topics with ChatGPTabstractLarge language models (LLMs) like ChatGPT can generate text for various prompts. With exceptional reasoning capabilities, ChatGPT (particularly the GPT-4 model) has achieved widespread adoption across many tasks - from creative writing to domain-specific inquiries, code generation, and more. This research analyzed the DevGPT dataset to determine common topics posed by developers interacting with ChatGPT. The DevGPT dataset comprises ChatGPT interactions from GitHub issues, pull requests and discussions. By employing a mixed-methods approach combining unsupervised semantic modeling and expert qualitative analysis we categorize the topics developers discuss when interacting with ChatGPT. Ertugrul Sagdic, Arda Bayram, Md. Rakibul Islam 0002 |
MSR | 3 |
| 2023 | Insights into Female Contributions in Open-Source ProjectsabstractThis paper presents a large quantitative study of the contributions of females compared to males in open-source projects. Female participation is found substantially low and females are found more engaged in non-coding work compared to men. The findings are statistically significant and are derived from an in-depth analysis of over 10 thousand developers’ contributions to more than 81 million different projects in the World of Code (WoC) infrastructure. The insights from this study are useful in addressing gender disparity in the field. Arifa I. Champa, Md. Fazle Rabbi, Minhaz Fahim Zibran, Md. Rakibul Islam 0002 |
MSR | 4 |
| 2018 | A comparison of software engineering domain specific sentiment analysis toolsabstractSentiment Analysis (SA) in software engineering (SE) text has drawn immense interests recently. The poor performance of general-purpose SA tools, when operated on SE text, has led to recent emergence of domain-specific SA tools especially designed for SE text. However, these domain-specific tools were tested on single dataset and their performances were compared mainly against general-purpose tools. Thus, two things remain unclear: (i) how well these tools really work on other datasets, and (ii) which tool to choose in which context. To address these concerns, we operate three recent domain-specific SA tools on three separate datasets. Using standard accuracy measurement metrics, we compute and compare their accuracies in the detection of sentiments in SE text. Md. Rakibul Islam 0002, Minhaz Fahim Zibran |
SANER | 1 |
| 2018 | SentiStrength-SE: Exploiting domain specificity for improved sentiment analysis in software engineering text
Md. Rakibul Islam 0002, Minhaz Fahim Zibran |
J. Syst. Softw. | 1 |
| 2017 | A Comparison of Dictionary Building Methods for Sentiment Analysis in Software Engineering TextabstractSentiment Analysis (SA) in Software Engineering (SE) texts suffers from low accuracies primarily due to the lack of an effective dictionary. The use of a domain-specific dictionary can improve the accuracy of SA in a particular domain. Building a domain dictionary is not a trivial task. The performance of lexical SA also varies based on the method applied to develop the dictionary. This paper includes a quantitative comparison of four dictionaries representing distinct dictionary building methods to identify which methods have higher/lower potential to perform well in constructing a domain dictionary for SA in SE texts. Md. Rakibul Islam 0002, Minhaz Fahim Zibran |
ESEM | 1 |
| 2017 | Security Vulnerabilities in Categories of Clones and Non-Cloned Code: An Empirical StudyabstractBackground: Software security has drawn immense importance in the recent years. While efforts are expected in minimizing security vulnerabilities in source code, the developers' practice of code cloning often causes multiplication of such vulnerabilities and program faults. Although previous studies examined the bug-proneness, stability, and changeability of clones against non-cloned code, the security aspects remained ignored. Aims: The objective of this work is to explore and understand the security vulnerabilities and their severity in different types of clones compared to non-clone code. Method: Using a state-of-the-art clone detector and two reputed security vulnerability detection tools, we detect clones and vulnerabilities in 8.7 million lines of code over 34 software systems. We perform a comparative study of the vulnerabilities identified in different types of clones and non-cloned code. The results are derived based on quantitative analyses with statistical significance. Results: Our study reveals that the security vulnerabilities found in code clones have higher severity of security risks compared to those in non-cloned code. However, the proportion (i.e., density) of vulnerabilities in clones and non-cloned code does not have any significant difference. Conclusion: The findings from this work add to our understanding of the characteristics and impacts of clones, which will be useful in clone-aware software development with improved software security. Md. Rakibul Islam 0002, Minhaz Fahim Zibran, Aayush Nagpal |
ESEM | 1 |
| 2017 | Leveraging automated sentiment analysis in software engineeringabstractAutomated sentiment analysis in software engineering textual artifacts has long been suffering from inaccuracies in those few tools available for the purpose. We conduct an in-depth qualitative study to identify the difficulties responsible for such low accuracy. Majority of the exposed difficulties are then carefully addressed in developing SentiStrength-SE, a tool for improved sentiment analysis especially designed for application in the software engineering domain. Using a benchmark dataset consisting of 5,600 manually annotated JIRA issue comments, we carry out both quantitative and qualitative evaluations of our tool. SentiStrength-SE achieves 73.85% precision and 85% recall, which are significantly higher than a state-of-the-art sentiment analysis tool we compare with. Md. Rakibul Islam 0002, Minhaz Fahim Zibran |
MSR | 1 |
| 2017 | Insights into continuous integration build failuresabstractContinuous integration is prevalently used in modern software engineering to build software systems automatically. Broken builds hinder developers' work and delay project progress. We must identify the factors causing build failures. This paper presents a large empirical study to identify the factors such as, complexity of a task, build strategy and contribution models (i.e., push and pull request), and projects level attributes (i.e., sizes of projects and teams), which potentially have impacts on the build results. We have studied 3.6 million builds over 1,090 open-source projects. The derived results add to our understanding of the role of those factors on build results, which can be used in minimizing build failures. Md. Rakibul Islam 0002, Minhaz Fahim Zibran |
MSR | 1 |
| 2016 | Towards understanding and exploiting developers' emotional variations in software engineeringabstractSoftware development is highly dependent on human efforts and collaborations, which are immensely affected by emotions. This paper presents a quantitative empirical study of the emotional variations in different types of development activities (e.g., bug-fixing tasks) and development periods (i.e., days and times), in addition to in-depth investigation of emotions' impacts on software artifacts (i.e., commit messages) and exploration of scopes for exploiting emotional variations in software engineering activities. We study emotions in more than 490 thousand commit comments across 50 open-source projects. The findings add to our understanding of the role of emotions in software development, and expose scopes for exploitation of emotional awareness in improved task assignments and collaborations. Md. Rakibul Islam 0002, Minhaz Fahim Zibran |
SERA | 1 |