Parvez Mahbub

dblp:336/4097 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-9164-5204ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Improved Detection and Diagnosis of Faults in Deep Neural Networks Using Hierarchical and Explainable Classification
abstract
Deep Neural Networks (DNN) have found numerous applications in various domains, including fraud detection, medical diagnosis, facial recognition, and autonomous driving. However, D NN - based systems often suffer from reliability issues due to their inherent complexity and the stochastic nature of their underlying models. Unfortunately, existing techniques to detect faults in DNN programs are either limited by the types of faults (e.g., hyperparameter or layer) they support or the kind of information (e.g., dynamic or static) they use. As a result, they might fall short of comprehensively detecting and diagnosing the faults. In this paper, we present DEFault (Detect and Explain Fault) - a novel technique to detect and diagnose faults in DNN programs. It first captures dynamic (i.e., runtime) features during model training and leverages a hierarchical classification approach to detect all major fault categories from the literature. Then, it captures static features (e.g., layer types) from DNN programs and leverages explainable AI methods (e.g., SHAP) to narrow down the root cause of the fault. We train and evaluate DEFault on a large, diverse dataset of ≈ 14.5K DNN programs and further validate our technique using a benchmark dataset of 52 real-life faulty DNN programs. Our approach achieves≈ 94% recall in detecting real-world faulty DNN programs and ≈ 63% recall in diagnosing the root causes of the faults, demonstrating 3.92%-11.54% higher performance than that of state-of-the-art techniques. Thus, DEFault has the potential to significantly improve the reliability of DNN programs by effectively detecting and diagnosing the faults.
Sigma Jahan, Mehil B. Shah, Parvez Mahbub, Mohammad Masudur Rahman 0001
ICSE3
2024 Predicting Line-Level Defects by Capturing Code Contexts with Hierarchical Transformers
abstract
Software defects consume 40% of the total budget in software development and cost the global economy billions of dollars every year. Unfortunately, despite the use of many software quality assurance (SQA) practices in software development (e.g., code review, continuous integration), defects may still exist in the official release of a software product. Therefore, prioritizing SQA efforts for the vulnerable areas of the codebase is essential to ensure the high quality of a software release. Predicting software defects at the line level could help prioritize the SQA effort but is a highly challenging task given that only ≈ 3% lines of a codebase could be defective. Existing works on line-level defect prediction often fall short and cannot fully leverage the line-level defect information. In this paper, we propose - Bugsplorer - a novel deep-learning technique for line-level defect prediction. It leverages a hierarchical structure of transformer models to represent two types of code elements: code tokens and code lines. Unlike the existing techniques that are optimized for file-level defect prediction, Bugsplorer is optimized for a line-level defect prediction objective. Our evaluation with five performance metrics shows that Bugsplorer has a promising capability of predicting defective lines with 26–72% better accuracy than that of the state-of-the-art technique. It can rank the first 20% defective lines within the top 1–3% suspicious lines. Thus, Bugsplorer has the potential to significantly reduce SQA costs by ranking defective lines higher.
Parvez Mahbub, Mohammad Masudur Rahman 0001
SANER1
2023 Explaining Software Bugs Leveraging Code Structures in Neural Machine Translation
abstract
Software bugs claim ≈ 50 % of development time and cost the global economy billions of dollars. Once a bug is reported, the assigned developer attempts to identify and understand the source code responsible for the bug and then corrects the code. Over the last five decades, there has been significant research on automatically finding or correcting software bugs. However, there has been little research on automatically explaining the bugs to the developers, which is essential but a highly challenging task. In this paper, we propose Bugsplainer, a transformer-based generative model, that generates natural language explanations for software bugs by learning from a large corpus of bug-fix commits. Bugsplainer can leverage structural information and buggy patterns from the source code to generate an explanation for a bug. Our evaluation using three performance metrics shows that Bugsplainer can generate understandable and good explanations according to Google's standard, and can outperform multiple baselines from the literature. We also conduct a developer study involving 20 participants where the explanations from Bugsplainer were found to be more accurate, more precise, more concise and more useful than the baselines.
Parvez Mahbub, Ohiduzzaman Shuvo, Mohammad Masudur Rahman 0001
ICSE1
2023 Bugsplainer: Leveraging Code Structures to Explain Software Bugs with Neural Machine Translation
abstract
Software bugs cost the global economy billions of dollars each year and take up ≈50% of the development time. Once a bug is reported, the assigned developer attempts to identify and understand the source code responsible for the bug and then corrects the code. Over the last five decades, there has been significant research on automatically finding or correcting software bugs. However, there has been little research on automatically explaining the bugs to the developers, which is essential but a highly challenging task. In this paper, we propose Bugsplainer, a novel web-based debugging solution that generates natural language explanations for software bugs by learning from a large corpus of bug-fix commits. Bugsplainer leverages code structures to reason about a bug and employs the fine-tuned version of a text generation model – CodeT5 – to generate the explanations.Tool video: https://youtu.be/xga-ScvULpk
Parvez Mahbub, Mohammad Masudur Rahman 0001, Ohiduzzaman Shuvo, Avinash Gopal
ICSME1
2023 Recommending Code Reviews Leveraging Code Changes with Structured Information Retrieval
abstract
Review comments are one of the main building blocks of modern code reviews. Manually writing code review comments could be time-consuming and technically challenging. Recently, an information retrieval (IR) based approach has been proposed to automatically recommend relevant code review comments for method-level code changes. However, this technique overlooks the structured items (e.g., class name, library information) from the source code and is applicable only for method-level changes. In this paper, we propose a novel technique for relevant review comments recommendation – RevCom – that leverages various code-level changes using structured information retrieval. RevCom uses different structured items from source code and can recommend relevant reviews for all types of changes (e.g., method-level and non-method-level). Our evaluation using three performance metrics show that RevCom outperforms both IR-based and DL-based baselines by up to 49.45% and 23.57% margins in BLEU score in recommending review comments. We find that RevCom can recommend review comments with an average BLEU score of ≈ 26.63%. According to Google’s AutoML Translation documentation, such a BLEU score indicates that the review comments can capture the original intent of the reviewers. All these findings suggest that RevCom can recommend relevant code reviews and has the potential to reduce the cognitive effort of human code reviewers.
Ohiduzzaman Shuvo, Parvez Mahbub, Mohammad Masudur Rahman 0001
ICSME2
2023 Defectors: A Large, Diverse Python Dataset for Defect Prediction
abstract
Defect prediction has been a popular research topic where machine learning (ML) and deep learning (DL) have found numerous applications. However, these ML/DL-based defect prediction models are often limited by the quality and size of their datasets. In this paper, we present Defectors, a large dataset for just-in-time and line-level defect prediction. Defectors consists of ≈ 213K source code files (≈ 93K defective and ≈ 120K defect- free) that span across 24 popular Python projects. These projects come from 18 different domains, including machine learning, automation, and internet-of-things. Such a scale and diversity make Defectors a suitable dataset for training ML/DL models, especially transformer models that require large and diverse datasets. We also foresee several application areas of our dataset including defect prediction and defect explanation.
Parvez Mahbub, Ohiduzzaman Shuvo, Mohammad Masudur Rahman 0001
MSR1