Saikat Mondal

dblp:192/3920 · DBLP profile ↗
← Back
21ranked-venue papers
14as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 19 · 12 first-author · 17 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Algorithm-Based Pipeline for Reliable and Intent-Preserving Code Translation with LLMs
abstract
Code translation, the automatic conversion of programs between languages, is a growing use case for Large Language Models (LLMs). However, direct one-shot translation often fails to preserve program intent, leading to errors in control flow, type handling, and I/O behavior. We propose an algorithm-based pipeline that introduces a language-neutral intermediate specification to capture these details before code generation. This study empirically evaluates the extent to which structured planning can improve translation accuracy and reliability relative to direct translation. We conduct an automated paired experiment – direct and algorithm-based to translate between Python and Java using five widely used LLMs on the Avatar and CodeNet datasets. For each combination (model, dataset, approach, and direction), we compile and execute the translated program and run the tests provided. We record compilation results, runtime behavior, timeouts (e.g., infinite loop), and test outcomes. We compute accuracy from these tests, counting a translation as correct only if it compiles, runs without exceptions or timeouts, and passes all tests. We then map every failed compile-time and runtime case to a unified, language-aware taxonomy and compare subtype frequencies between the direct and algorithm-based approaches. Overall, the Algorithm-based approach increases micro-average accuracy from 67.7% to 78.5% (↑ 10.8%). It eliminates lexical and token errors by 100%, reduces incomplete constructs by 72.7%, and structural and declaration issues by 61.1%. It also substantially lowers runtime dependency and entry-point failures by 78.4%. These results demonstrate that algorithm-based pipelines enable more reliable, intent-preserving code translation, providing a foundation for robust multilingual programming assistants.
Shahriar Rumi Dipto, Saikat Mondal, Chanchal Kumar Roy
ICPC2
2026 Why Are AI Agent-Involved Pull Requests (Fix-Related) Remain Unmerged? An Empirical Study
abstract
Autonomous coding agents (e.g., OpenAI Codex, Devin, GitHub Copilot) are increasingly used to generate fix-related pull requests (PRs) in real-world software repositories. However, their practical effectiveness depends on whether project maintainers accept and merge these contributions. In this paper, we present an empirical study of AI agent–involved fix-related PRs, examining both their integration outcomes, latency, and the factors that hinder successful merging. We first analyze 8,106 fix-related PRs authored by five widely used AI coding agents from the AIDEV-POP dataset to quantify the proportions of PRs that are merged, closed without merging, or remain open. We then conduct a manual analysis of a statistically significant sample of 326 closed but unmerged PRs, spending approximately 100 person-hours to construct a structured catalog of 12 failure reasons. Our results indicate that test case failures and prior resolution of the same issues by other PRs are the most common causes of non-integration, whereas build or deployment failures are comparatively rare. Overall, our findings expose key limitations of current AI coding agents in real-world settings and highlight directions for their further improvement and for more effective human-AI collaboration in software maintenance.
Khairul Alam, Saikat Mondal, Banani Roy
MSR2
2026 Automatic assistance to mitigate rollback inconsistencies in collaborative edits
Saikat Mondal, Gias Uddin 0001, Chanchal Kumar Roy
Autom. Softw. Eng.1
2025 From Questions to Insights: Exploring XAI Challenges Reported on Stack Overflow Questions
abstract
The lack of interpretability is a major barrier that limits the practical usage of AI models. Several eXplainable AI (XAI) techniques (e.g., SHAP, LIME) have been employed to interpret these models’ performance. However, users often face challenges when leveraging these techniques in real-world scenarios and thus submit questions in technical Q&A forums like Stack Overflow (SO) to resolve these challenges. We conducted an exploratory study to expose these challenges, their severity, and features that can make XAI techniques more accessible and easier to use. Our contributions to this study are fourfold. First, we manually analyzed 663 SO questions that discussed challenges related to XAI techniques. Our careful investigation produced a catalog of seven challenges (e.g., disagreement issues). We then analyzed their prevalence and found that model integration and disagreement issues emerged as the most prevalent challenges. Second, we attempt to estimate the severity of each XAI challenge by determining the correlation between challenge types and answer metadata (e.g., the presence of accepted answers). Our analysis suggests that model integration issues is the most severe challenge. Third, we attempt to perceive the severity of these challenges based on practitioners’ ability to use XAI techniques effectively in their work. Practitioners’ responses suggest that disagreement issues most severely affect the use of XAI techniques. Fourth, we seek agreement from practitioners on improvements or features that could make XAI techniques more accessible and user-friendly. The majority of them suggest consistency in explanations and simplified integration. Our study findings might (a) help to enhance the accessibility and usability of XAI and (b) act as the initial benchmark that can inspire future research.
Saumendu Roy, Saikat Mondal, Banani Roy, Chanchal Kumar Roy
EASE2
2025 Does Editing Improve Answer Quality on Stack Overflow? A Data-Driven Investigation
abstract
High-quality answers in technical Q&A platforms like Stack Overflow (SO) are crucial as they directly influence software development practices. Poor-quality answers can introduce inefficiencies, bugs, and security vulnerabilities, and thus increase maintenance costs and technical debt in production software. To improve content quality, SO allows collaborative editing, where users revise answers to enhance clarity, correctness, and formatting. Several studies have examined rejected edits and identified the causes of rejection. However, prior research has not systematically assessed whether accepted edits enhance key quality dimensions. While one study investigated the impact of edits on$\mathrm{C} / \mathrm{C}++$vulnerabilities, broader quality aspects remain unexplored. In this study, we analyze 94,994 Python-related answers that have at least one accepted edit to determine whether edits improve (1) semantic relevance, (2) code usability, (3) code complexity, (4) security vulnerabilities, (5) code optimization, and (6) readability. Our findings show both positive and negative effects of edits. While 53.3% of edits improve how well answers match questions, 38.1% make them less relevant. Some previously broken code (9%) becomes executable, yet working code (14.7%) turns non-parsable after edits. Many edits increase complexity (32.3%), making code harder to maintain. Instead of fixing security issues, 20.5% of edits introduce additional issues. Even though 51.0% of edits optimize performance, execution time still increases overall. Readability also suffers, as 49.7% of edits make code harder to read. This study highlights the inconsistencies in editing outcomes and provides insights into how edits impact software maintainability, security, and efficiency that might caution users and moderators and help future improvements in collaborative editing systems.
Saikat Mondal, Chanchal Kumar Roy
ICSME1
2025 Why Do Developers Engage with ChatGPT in Issue-Tracker? Investigating Usage and Reliance on ChatGPT-Generated Code
abstract
Large language models (LLMs) like ChatGPT have shown the potential to assist developers with coding and debugging tasks. However, their role in collaborative issue resolution is underexplored. In this study, we analyzed 1,152 Developer-ChatGPT conversations across 1,012 issues in GitHub to examine the diverse usage of ChatGPT and reliance on its generated code. Our contributions are fourfold. First, we manually analyzed 289 conversations to understand ChatGPT's usage in the GitHub Issues. Our analysis revealed that ChatGPT is primarily utilized for ideation, whereas its usage for validation (e.g., code documentation accuracy) is minimal. Second, we applied BERTopic modeling to identify key areas of engagement on the entire dataset. We found that backend issues (e.g., API management) dominate conversations, while testing is surprisingly less covered. Third, we utilized the CPD clone detection tool to check if the code generated by ChatGPT was used to address issues. Our findings revealed that ChatGPT-generated code was used as-is to resolve only 5.83% of the issues. Fourth, we estimated sentiment using a RoBERTa-based sentiment analysis model to determine developers' satisfaction with different usages and engagement areas. We found positive sentiment (i.e., high satisfaction) about using ChatGPT for refactoring and addressing data analytics (e.g., categorizing table data) issues. On the contrary, we observed negative sentiment when using ChatGPT to debug issues and address automation tasks (e.g., GUI interactions). Our findings show the unmet needs and growing dissatisfaction among developers. Researchers and ChatGPT developers should focus on developing task-specific solutions that help resolve diverse issues, improving user satisfaction and problem-solving efficiency in software development.
Joy Krishan Das, Saikat Mondal, Chanchal Kumar Roy
SANER2
2024 Investigating the Utility of ChatGPT in the Issue Tracking System: An Exploratory Study
abstract
Issue tracking systems serve as the primary tool for incorporating external users and customizing a software project to meet the users' requirements. However, the limited number of contributors and the challenge of identifying the best approach for each issue often impede effective resolution. Recently, an increasing number of developers are turning to AI tools like ChatGPT to enhance problem-solving efficiency. While previous studies have demonstrated the potential of ChatGPT in areas such as automatic program repair, debugging, and code generation, there is a lack of study on how developers explicitly utilize ChatGPT to resolve issues in their tracking system. Hence, this study aims to examine the interaction between ChatGPT and developers to analyze their prevalent activities and provide a resolution. In addition, we assess the code reliability by confirming if the code produced by ChatGPT was integrated into the project's codebase using the clone detection tool NiCad. Our investigation reveals that developers mainly use ChatGPT for brainstorming solutions but often opt to write their code instead of using ChatGPT-generated code, possibly due to concerns over the generation of "hallucinated" code, as highlighted in the literature.
Joy Krishan Das, Saikat Mondal, Chanchal Kumar Roy
MSR2
2024 Enhancing User Interaction in ChatGPT: Characterizing and Consolidating Multiple Prompts for Issue Resolution
abstract
Prompt design plays a crucial role in shaping the efficacy of ChatGPT, influencing the model's ability to extract contextually accurate responses. Thus, optimal prompt construction is essential for maximizing the utility and performance of ChatGPT. However, sub-optimal prompt design may necessitate iterative refinement, as imprecise or ambiguous instructions can lead to undesired responses from ChatGPT. Existing studies explore several prompt patterns and strategies to improve the relevance of responses generated by ChatGPT. However, the exploration of constraints that necessitate the submission of multiple prompts is still an unmet attempt. In this study, our contributions are twofold. First, we attempt to uncover gaps in prompt design that demand multiple iterations. In particular, we manually analyze 686 prompts that were submitted to resolve issues related to Java and Python programming languages and identify eleven prompt design gaps (e.g., missing specifications). Such gap exploration can enhance the efficacy of single prompts in ChatGPT. Second, we attempt to reproduce the ChatGPT response by consolidating multiple prompts into a single one. We can completely consolidate prompts with four gaps (e.g., missing context) and partially consolidate prompts with three gaps (e.g., additional functionality). Such an effort provides concrete evidence to users to design more optimal prompts mitigating these gaps. Our study findings and evidence can - (a) save users time, (b) reduce costs, and (c) increase user satisfaction.
Saikat Mondal, Suborno Deb Bappon, Chanchal Kumar Roy
MSR1
2024 AUTOGENICS: Automated Generation of Context-Aware Inline Comments for Code Snippets on Programming Q&A Sites Using LLM
abstract
Inline comments in the source code facilitate easy comprehension, reusability, and enhanced readability. However, code snippets in answers on Q&A sites like Stack Overflow (SO) often lack comments because answerers volunteer their time and often skip comments or explanations due to time constraints. Existing studies show that these online code examples are difficult to read and understand, making it difficult for developers (espe-cially novices) to use them correctly and leading to misuse. Given these challenges, we introduced AUTOGENICS, a tool designed to integrate with SO to generate effective inline comments for code snippets in SO answers exploiting large language models (LLMs). Our contributions are threefold. First, we randomly select 400 answer code snippets (200 Python + 200 Java) from SO and gener-ate inline comments for them using LLMs (e.g., Gemini). We then manually evaluate these comments' effectiveness using four key metrics: accuracy, adequacy, conciseness, and usefulness. Overall, LLMs demonstrate promising effectiveness in generating inline comments for SO answer code snippets. Second, we surveyed 14 active SO users to perceive the effectiveness of these inline comments. The survey results are consistent with our previous manual evaluation. However, according to our evaluation, LLMs-generated comments are less effective for shorter code snippets and sometimes produce noisy comments. Third, to address the gaps, we introduced AUTOGENICS that extracts additional context from question texts and generates context-aware inline comments. It also optimizes comments by removing noise (e.g., comments in import statements and variable declarations). We evaluate the effectiveness of AUTOGENICS-generated comments using the same four metrics that outperform those of standard LLMs. AUTOGENICS might (a) enhance code comprehension with context-aware inline comments, (b) save time, and improve developers' ability to learn and reuse code more accurately.
Suborno Deb Bappon, Saikat Mondal, Banani Roy
SCAM2
2024 Can We Identify Stack Overflow Questions Requiring Code Snippets? Investigating the Cause & Effect of Missing Code Snippets
abstract
On the Stack Overflow (SO) Q&A site, users often request solutions to their code-related problems (e.g., errors, unexpected behavior). Unfortunately, they often miss required code snippets during their question submission. Such a practice could prevent their questions from getting prompt and appropriate answers. In this study, we conduct an empirical study investigating the cause & effect of missing code snippets in SO questions whenever required. In this paper, our contributions are threefold. First, we analyze how the presence or absence of required code snippets in SO questions affects the correlation between question types (missed code, included code after requests & had code snippets during submission) and corresponding answer meta-data, such as the presence of an accepted answer. According to our analysis, the chance of getting accepted answers is three times higher for questions that include required code snippets during their question submission than those that missed the code. We also investigate the confounding factors (e.g., user reputation) that can affect questions receiving answers besides the presence or absence of required code snippets. We found that such factors do not hurt the correlation between the presence or absence of required code snippets and answer meta-data. Second, we surveyed 64 practitioners to understand why users miss necessary code snippets. About 60% of them agree that users are unaware of whether their questions require any code snippets. Third, we thus extract four text-based features (e.g., keywords, POS-based patterns) and build six Machine Learning (ML) models to identify the questions that need code snippets. Our models can predict the target questions with 86.5 % precision, 90.8 % recall, 85.3 % F1-score, and 85.2 % overall accuracy, which are highly promising. Our work has the potential to ($a$) save significant time in programming question-answering and (b) improve the quality of the valuable knowledge base by decreasing unanswered and unresolved questions.
Saikat Mondal, Mohammad Masudur Rahman 0001, Chanchal Kumar Roy
SANER1
2024 Reproducibility of issues reported in stack overflow questions: Challenges, impact & estimation
Saikat Mondal, Banani Roy
J. Syst. Softw.1
2023 Investigating Technology Usage Span by Analyzing Users' Q&A Traces in Stack Overflow
abstract
Choosing an appropriate software development technology (e.g., programming language) is challenging due to the proliferation of diverse options. The selection of inappropriate technologies for development may have a far-reaching effect on software developers' career growth. Switching to a different technology after working with one may lead to a complex learning curve and, thus, be more challenging. Therefore, it is crucial for software developers to find technologies that have a high usage span. Intuitively, the usage span of a technology can be deter-mined by the time span developers have used that technology. Existing literature focuses on the technology landscape to explore the complex and implicit dependencies among technologies but lacks formal studies to draw insights about their usage span. This paper investigates the technology usage span by analyzing the question and answering (Q&A) traces of Stack Overflow (SO), the largest technical Q&A website available to date. In particular, we analyze 6.7 million Q&A traces posted by about 97K active SO users and see what technologies have appeared in their questions or answers over 15 years. According to our analysis, C# and Java programming languages have a high usage span, followed by JavaScript. Besides, developers used the. NET framework, iO$S$& Windows Operating Systems (OS), and SQL query language for a long time (on average). Our study also exposes the emerging (i.e., newly growing) technologies. For example, usages of technologies such as SwiftUI,. NET-6.0, Visual Studio 2022, and Blazor WebAssembly framework are increasing. The findings from our study can assist novice developers, startup software industries, and software users in determining appropriate technologies. This also establishes an initial benchmark for future investigation on the use span of software technologies.
Saikat Mondal, Debajyoti Mondal, Chanchal Kumar Roy
APSEC1
2023 Do Subjectivity and Objectivity Always Agreeƒ A Case Study with Stack Overflow Questions
abstract
In Stack Overflow (SO), the quality of posts (i.e., questions and answers) is subjectively evaluated by users through a voting mechanism. The net votes (upvotes − downvotes) obtained by a post are often considered an approximation of its quality. However, about half of the questions that received working solutions got more downvotes than upvotes. Furthermore, about 18% of the accepted answers (i.e., verified solutions) also do not score the maximum votes. All these counter-intuitive findings cast doubts on the reliability of the evaluation mechanism employed at SO. Moreover, many users raise concerns against the evaluation, especially downvotes to their posts. Therefore, rigorous verification of the subjective evaluation is highly warranted to ensure a non-biased and reliable quality assessment mechanism. In this paper, we compare the subjective assessment of questions with their objective assessment using 2.5 million questions and ten text analysis metrics. According to our investigation, four objective metrics agree with the subjective evaluation, two do not agree, one either agrees or disagrees, and the remaining three neither agree nor disagree with the subjective evaluation. We then develop machine learning models to classify the promoted and discouraged questions. Our models outperform the state-of-the-art models with a maximum of about 76%–87% accuracy.
Saikat Mondal, Mohammad Masudur Rahman 0001, Chanchal Kumar Roy
MSR1
2023 Automatic prediction of rejected edits in Stack Overflow
Saikat Mondal, Gias Uddin 0001, Chanchal Kumar Roy
Empir. Softw. Eng.1
2022 Why Don't XAI Techniques Agree? Characterizing the Disagreements Between Post-hoc Explanations of Defect Predictions
abstract
Machine Learning (ML) based defect prediction models can be used to improve the reliability and overall quality of software systems. However, such defect predictors might not be deployed in real applications due to the lack of transparency. Thus, recently, application of several post-hoc explanation methods (e.g., LIME and SHAP) have gained popularity. These explanation methods can offer insight by ranking features based on their importance in black box decisions. The explainability of ML techniques is reasonably novel in the Software Engineering community. However, it is still unclear whether such explainability methods genuinely help practitioners make better decisions regarding software maintenance. Recent user studies show that data scientists usually utilize multiple post-hoc explainers to understand a single model decision because of the lack of ground truth. Such a scenario causes disagreement between explainability methods and impedes drawing a conclusion. Therefore, our study first investigates three disagreement metrics between LIME and SHAP explanations of 10 defect-predictors, and exposes that disagreements regarding the rankings of feature importance are most frequent. Our findings lead us to propose a method of aggregating LIME and SHAP explanations that puts less emphasis on these disagreements while highlighting the aspect on which explanations agree.
Saumendu Roy, Gabriel Laberge, Banani Roy, Foutse Khomh, Amin Nikanjam, Saikat Mondal
ICSME6
2022 The reproducibility of programming-related issues in Stack Overflow questions
Saikat Mondal, Mohammad Masudur Rahman 0001, Chanchal Kumar Roy, Kevin A. Schneider
Empir. Softw. Eng.1
2021 Rollback Edit Inconsistencies in Developer Forum
abstract
The success of developer forums like Stack Overflow (SO) depends on the participation of users and the quality of shared knowledge. SO allows its users to suggest edits to improve the quality of the posts (i.e., questions and answers). Such posts can be rolled back to an earlier version when the current version of the post with the suggested edit does not satisfy the user. However, subjectivity bias in deciding either an edit is satisfactory or not could introduce inconsistencies in the rollback edits. For example, while a user may accept the formatting of a method name (e.g., getActivity()) as a code term, another user may reject it. Such bias in rollback edits could be detrimental and demotivating to the users whose suggested edits were rolled back. This problem is compounded due to the absence of specific guidelines and tools to support consistency across users on their rollback actions. To mitigate this problem, we investigate the inconsistencies in the rollback editing process of SO and make three contributions. First, we identify eight inconsistency types in rollback edits through a qualitative analysis of 777 rollback edits in 382 questions and 395 answers. Second, we determine the impact of the eight rollback inconsistencies by surveying 44 software developers. More than 80% of the study participants find our produced catalogue of rollback inconsistencies to be detrimental to the post quality. Third, we develop a suite of algorithms to detect the eight rollback inconsistencies. The algorithms offer more than 95% accuracy and thus can be used to automatically but reliably inform users in SO of the prevalence of inconsistencies in their suggested edits and rollback actions.
Saikat Mondal, Gias Uddin 0001, Chanchal Kumar Roy
MSR1
2020 Automatic Identification of Rollback Edit with Reasons in Stack Overflow Q&A Site
abstract
Crowd-sourced developer forums, such as Stack Overflow (SO), rely on edits from users to improve the quality of the shared knowledge. Unfortunately, suggested edits in SO are frequently rejected by rollbacks due to undesired edits or violation of editing guidelines. Such rollbacks could frustrate and demotivate users to provide future suggestions. We thus need to warn a user of a potential rollback so that he can improve the suggested edit and thus increase its likelihood of acceptance. This study proposes to help users with an automated machine learning classification model that can warn them of potential rollbacks to their suggested edits. We present the conceptual design of EditEx, an online tool that can guide SO users during their editing by highlighting the potential causes of rollback. We offer details of an empirical study to assess the accuracy of the classifiers and a user study to evaluate the effectiveness of EditEx.
Saikat Mondal, Gias Uddin 0001, Chanchal Kumar Roy
ICSME1
2019 Can issues reported at stack overflow questions be reproduced?: an exploratory study
abstract
Software developers often look for solutions to their code level problems at Stack Overflow. Hence, they frequently submit their questions with sample code segments and issue descriptions. Unfortunately, it is not always possible to reproduce their reported issues from such code segments. This phenomenon might prevent their questions from getting prompt and appropriate solutions. In this paper, we report an exploratory study on the reproducibility of the issues discussed in 400 questions of Stack Overflow. In particular, we parse, compile, execute and even carefully examine the code segments from these questions, spent a total of 200 man hours, and then attempt to reproduce their programming issues. The outcomes of our study are two-fold. First, we find that 68% of the code segments require minor and major modifications in order to reproduce the issues reported by the developers. On the contrary, 22% code segments completely fail to reproduce the issues. We also carefully investigate why these issues could not be reproduced and then provide evidence-based guidelines for writing effective code examples for Stack Overflow questions. Second, we investigate the correlation between issue reproducibility status (of questions) and corresponding answer meta-data such as the presence of an accepted answer. According to our analysis, a question with reproducible issues has at least three times higher chance of receiving an accepted answer than the question with irreproducible issues.
Saikat Mondal, Mohammad Masudur Rahman 0001, Chanchal Kumar Roy
MSR1
2019 Blockchain Inspired RFID-Based Information Architecture for Food Supply Chain
abstract
In this paper, we propose a blockchain inspired Internet-of-Things architecture for creating a transparent food supply chain. The architecture uses a proof-of-object-based authentication protocol, which is analogous to the cryptocurrency's proof-of-work protocol. The complete architecture was realized by integrating a radio frequency identification (RFID)-based sensor at the physical layer and blockchain at the cyber layer. The RFID provides a unique identity of the product and the sensor data, which helps in real time quality monitoring. For this purpose, a small feature size 900-MHz RFID coupled sensor was fabricated and demonstrated for real time sensor data acquisition. The blockchain architecture aids in creating a tamper-proof digital database of the food packages at each instance. A detailed security analysis was performed to investigate the vulnerability of the proposed architecture under different types of cyber attacks.
Saikat Mondal, Kanishka P. Wijewardena, Saranraj Karuppuswami, Nitya Kriti, Premjeet Chahal
IEEE Internet Things J.1
2017 Modeling and Crosstalk Evaluation of 3-D TSV-Based Inductor With Ground TSV Shielding
abstract
In this paper, we present a novel through-silicon-via (TSV)-based 3-D inductor structure with ground TSV shielding for better noise performance. In addition, a circuit model is proposed for the inductor, which can reduce the simulation time over finite-element-based 3-D full-wave simulation. Rigorous 3-D full-wave simulation is performed up to 10 GHz to validate the circuit model. The ground TSV-based 3-D inductor is found to be resilient to TSV-TSV crosstalk noise compared with conventional 3-D inductors. The simulation results revealed that more than -33 dB of isolation can be achieved at 2 GHz between the 3-D inductor and the noise probe.
Saikat Mondal, Sang-Bock Cho, Bruce C. Kim
IEEE Trans. Very Large Scale Integr. Syst.1