VLDB 2026 Research / reviewers in the wild / expert
Shamse Tasnim Cynthia
dblp:252/2943
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0001-9529-0132ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Are We All Using Agents the Same Way? An Empirical Study of Core and Peripheral Developers' Use of Coding AgentsabstractAutonomous AI agents are transforming software development and redefining how developers collaborate with AI. Prior research shows that the adoption and use of AI-powered tools differ between core and peripheral developers. However, it remains unclear how this dynamic unfolds in the emerging era of autonomous coding agents. In this paper, we present the first empirical study of 9,427 agentic PRs, examining how core and peripheral developers use, review, modify, and verify agent-generated contributions prior to acceptance. Through a mix of qualitative and quantitative analysis, we make four key contributions. First, a subset of peripheral developers use agents more often, delegating tasks evenly across bug fixing, feature addition, documentation, and testing. In contrast, core developers focus more on documentation and testing, yet their agentic PRs are frequently merged into the main/master branch. Second, core developers engage slightly more in review discussions than peripheral developers, and both groups focus on evolvability issues. Third, agentic PRs are less likely to be modified, but when they are, both groups commonly perform refactoring. Finally, peripheral developers are more likely to merge without running CI checks, whereas core developers more consistently require passing verification before acceptance. Our analysis offers a comprehensive view of how developer experience shapes integration offer insights for both peripheral and core developers on how to effectively collaborate with coding agents. Shamse Tasnim Cynthia, Joy Krishan Das, Banani Roy |
MSR | 1 |
| 2026 | Beyond Bug Fixes: An Empirical Investigation of Post-Merge Code Quality Issues in Agent-Generated Pull RequestsabstractThe increasing adoption of AI coding agents has increased the number of agent-generated pull requests (PRs) merged with little or no human intervention. Although such PRs promise productivity gains, their post-merge code quality remains underexplored, as prior work has largely relied on benchmarks and controlled tasks rather than large-scale post-merge analyses. To address this gap, we analyze 1,210 merged agent-generated bug-fix PRs from Python repositories in the AIDev dataset. Using SonarQube, we perform a differential analysis between base and merged commits to identify code quality issues newly introduced by PR changes. We examine issue frequency, density, severity, and rule-level prevalence across five agents. Our results show that apparent differences in raw issue counts across agents largely disappear after normalizing by code churn, indicating that higher issue counts are primarily driven by larger PRs. Across all agents, code smells dominate, particularly at critical and major severities, while bugs are less frequent but often severe. Overall, our findings show that merge success does not reliably reflect post-merge code quality, highlighting the need for systematic quality checks for agent-generated bug-fix PRs. Shamse Tasnim Cynthia, Al Muttakin, Banani Roy |
MSR | 1 |
| 2025 | Identification and Optimization of Redundant Code Using Large Language ModelsabstractRedundant code is a persistent challenge in software development that makes systems harder to maintain, scale, and update. It adds unnecessary complexity, hinders bug fixes, and increases technical debt. Despite their impact, removing redundant code manually is risky and error-prone, often introducing new bugs or missing dependencies. While studies highlight the prevalence and negative impact of redundant code, little focus has been given to Artificial Intelligence (AI) system codebases and the common patterns that cause redundancy. Additionally, the reasons behind developers unintentionally introducing redundant code remain largely unexplored. This research addresses these gaps by leveraging large language models (LLMs) to automatically detect and optimize redundant code in AI projects. Our research aims to identify recurring patterns of redundancy and analyze their underlying causes, such as outdated practices or insufficient awareness of best coding principles. Additionally, we plan to propose an LLM agent that will facilitate the detection and refactoring of redundancies on a large scale while preserving original functionality. This work advances the application of AI in identifying and optimizing redundant code, ultimately helping developers maintain cleaner, more readable, and scalable codebases. Shamse Tasnim Cynthia |
CAIN | 1 |
| 2025 | How do Community Smells Influence Self-Admitted Technical Debt in Machine Learning Projects?abstractBackground: Community smells reflect poor organizational practices that often lead to socio-technical issues and the accumulation of Self-Admitted Technical Debt (SATD). While prior studies have explored these problems in general software systems, their interplay in machine learning (ML)-based projects remains largely under-examined. Aims: In this study, we aim to investigate the prevalence of community smells and their relationship with SATD in open-source ML projects, analyzing data at the release level. Methods: We analyzed$\mathbf{1 5 5 ~ M L}$-based systems across multiple releases to examine the prevalence of ten community smell types. Then we detected SATD at the release level and applied statistical analysis to examine its correlation with community smells. Next, we considered the six identified types of SATD to determine which community smells are the most associated with each debt category. Finally, we analyzed how the community smells and SATD evolve over the releases, uncovering project size-dependent trends and shared trajectories. Results: Community smells are found to be widespread, exhibiting distinct distribution patterns across small, medium and large projects. Certain smells, such as Radio Silence and Organizational Silos, are strongly correlated with higher SATD occurrences, while authority- and communication-related smells often co-occur with persistent code and design debt. Temporal analysis revealed shared evolutionary trajectories of smells and SATD, influenced by project size. Conclusion: Our findings emphasize the importance of early detection and mitigation of socio-technical issues to maintain the long-term quality and sustainability of ML-based systems. Shamse Tasnim Cynthia, Nuri Almarimi, Banani Roy |
ESEM | 1 |
| 2025 | Feature transformation for improved software bug detection and commit classificationabstractTesting and debugging software to fix bugs is considered one of the most important stages of the software life cycle. Many studies have investigated ways to predict bugs in software artifacts using machine learning techniques. It is important to consider the explanatory aspects of such models for reliable prediction. In this paper, we show how feature transformation can significantly improve prediction accuracy and provide insight into the inner workings of bug prediction models. We propose a new approach for bug prediction that first extracts the features, then finds a weighted transformation of these features using a genetic algorithm that best separates bugs from non-bugs when plotted in a low-dimensional space, and finally, trains predictive models using the transformed dataset. In our experiment using the proposed feature transformation, the traditional machine learning and deep learning classifiers achieved an average improvement of 4.25% and 9.6% in recall values for bug classification over 8 software systems compared to the models built on original data. We also examined the generalizability of our concept for multiclass classification tasks such as commit classification in software systems and found modest improvements in F1-scores (sometimes up to 3%) for traditional machine learning models and 4% with deep learning models. • Feature transformation techniques applied to bug detection in software systems. • Genetic algorithm based transformation and t-SNE based clustering in low dimensions. • Improved explainability and accuracy of machine learning based bug detection models. • Applicable to deep learning models and generalizable to commit classification. Sakib Mostafa, Shamse Tasnim Cynthia, Banani Roy, Debajyoti Mondal |
J. Syst. Softw. | 2 |