VLDB 2026 Research / reviewers in the wild / expert
Sushant Kumar Pandey
dblp:253/1149
· DBLP profile ↗
15ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0003-1882-2435ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Software engineering, systems software and programming languages · 8 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Train-Test Data Leakage in Automotive Image DatasetsabstractReliable evaluation of machine learning (ML)-enabled perception systems for intelligent vehicles critically depends on the integrity of training and test datasets. A major risk arises when near-duplicate or visually similar images appear across subsets, leading to inflated performance estimates. This study systematically quantifies train-test similarity in six widely used automotive datasets - KITTI, ZOD, BDD100k, ONCE, Cirrus, and SODA10M - in their default splits. We employ perceptual hashing (pHash) and deep feature embeddings to measure image-level redundancy. Results show 4,513 pairs of KITTI and 562 of ZOD images were almost identical, corresponding to 25% and 5% of their test images, respectively. The other examined datasets contain only at most 3 pairs (almost 0%) of almost identical images in their existing train-test splits. These findings underscore the need for similarity analysis during dataset preparation, particularly for video-based collections with strong spatio-temporal dependencies. By exposing dataset-specific risks of data leakage in popular datasets, this study contributes practical insights for both dataset curators and ML practitioners. These insights are valuable when using public benchmark datasets in safety-critical domains such as autonomous driving (AD). Md. Abu Ahammed Babu, Miroslaw Staron, András Bálint, Darko Durisic, Sushant Kumar Pandey |
IV | 5 |
| 2025 | How Well Small Language Models Can Be Adapted for Software Maintenance and Refactoring Tasks
Gabija Asvydyte, Sushant Kumar Pandey, Sivajeet Chand |
PROFES | 2 |
| 2025 | D-LeDe: A Data Leakage Detection Method for Automotive Perception SystemsabstractData leakage is a very common problem that is often overlooked during splitting data into train and test sets before training any ML/DL model. The model performance gets artificially inflated with the presence of data leakage during the evaluation phase which often leads the model to erroneous prediction on real-time deployment. However, detecting the presence of such leakage is challenging, particularly in the object detection context of perception systems where the model needs to be supplied with image data for training. In this study, we conduct a computational experiment to develop a method for detecting data leakage. We then conducted an initial evaluation of the method as a first step on a public dataset, “Kitti”, which is a popular and widely accepted benchmark dataset in the automotive domain. The evaluation results show that our proposed D-LeDe method are able to successfully detect potential data leakage caused by image similarity. A further validation was also provided to justify the evaluation outcome by conducting pair-wise image similarity analysis using perceptual hash (pHash) distance. Md. Abu Ahammed Babu, Sushant Kumar Pandey, Darko Durisic, Ashok Chaitanya Koppisetty, Miroslaw Staron |
VEHITS | 2 |
| 2025 | Design pattern recognition: a study of large language modelsabstractAbstract Context As Software Engineering (SE) practices evolve due to extensive increases in software size and complexity, the importance of tools to analyze and understand source code grows significantly. Objective This study aims to evaluate the abilities of Large Language Models (LLMs) in identifying DPs in source code, which can facilitate the development of better Design Pattern Recognition (DPR) tools. We compare the effectiveness of different LLMs in capturing semantic information relevant to the DPR task. Methods We studied Gang of Four (GoF) DPs from the P-MARt repository of curated Java projects. State-of-the-art language models, including Code2Vec, CodeBERT, CodeGPT, CodeT5, and RoBERTa, are used to generate embeddings from source code. These embeddings are then used for DPR via a k-nearest neighbors prediction. Precision, recall, and F1-score metrics are computed to evaluate performance. Results RoBERTa is the top performer, followed by CodeGPT and CodeBERT, which showed mean F1 Scores of 0.91, 0.79, and 0.77, respectively. The results show that LLMs without explicit pre-training can effectively store semantics and syntactic information, which can be used in building better DPR tools. Conclusion The performance of LLMs in DPR is comparable to existing state-of-the-art methods but with less effort in identifying pattern-specific rules and pre-training. Factors influencing prediction performance in Java files/programs are analyzed. These findings can advance software engineering practices and show the importance and abilities of LLMs for effective DPR in source code. Sushant Kumar Pandey, Sivajeet Chand, Jennifer Horkoff, Miroslaw Staron, Miroslaw Ochodek, Darko Durisic |
Empir. Softw. Eng. | 1 |
| 2023 | Cross-Project setting using Deep learning Architectures in Just-In-Time Software Fault Prediction: An InvestigationabstractThe prediction of whether a software change is fault-inducing or not in the software system using various learning methods, the study concerned in Just-In-Time Software Fault Prediction (JIT-SFP). Building such predicting model requires adequate training data. However, there needs to be more training data at the beginning of the software system. Cross-Project (CP) setting can subjugate this challenge by employing data from different software projects. It can achieve similar predictive performance to Within-Project (WP) fault prediction. It is still being determined to what level the CP training data can be useful in such a situation. Furthermore, it also needs to be discovered whether CP data are helpful in the initial phase of fault detection, and when there is an inadequate WP train set, CP could be beneficial to extend. This article deals with such investigations in real software projects. We proposed a new method by levering a deep belief network and long short-term memory called JITCP-Predictor. Out of ten, the proposed model significantly outperforms every ten project benchmark methods, and it is superior from 10.63% to 136.36% and 7.04% to 35.71% in terms of MCC and F-Measure, respectively. The mean values of MCC and F-Measure produced by JITCP-Predictor are 0.52 ± 0.021 and 0.76 ± 0.76, respectively. We also found that the proposed model is more suitable for large and moderate-size projects. The proposed model avoids class imbalance and overfitting problems and takes reasonable training costs. Sushant Kumar Pandey, Anil Kumar Tripathi |
AST | 1 |
| 2023 | TransDPR: Design Pattern Recognition Using Programming Language ModelsabstractCurrent Design Pattern Recognition (DPR) methods have limitations, such as the reliance on semantic information, limited recognition of novel or modified pattern versions, and other factors. We present an introductory DPR technique by using a Programming Language Model (PLM) called TransDPR, which utilizes a Facebook pre-trained model (TransCoder), which is a Cross-lingual programming Language Model (XLM) based on a transformer architecture. We leverage an n-dimensional vector representation of programs and apply logistic regression to learn design patterns (DPs). Our approach utilizes the GitHub repository to collect singleton and prototype DP programs written in$C$++ source code. Our results indicate that TransDPR achieves 90% accuracy and an F1-score of 0.88 on open-source projects. We evaluate the proposed model on two developed modules from Volvo Cars and invite the original developers to validate the prediction results. Sushant Kumar Pandey, Miroslaw Staron, Jennifer Horkoff, Miroslaw Ochodek, Nicholas Mucci, Darko Durisic |
ESEM | 1 |
| 2023 | Defect Backlog Size Prediction for Open-Source Projects with the Autoregressive Moving Average and Exponential Smoothing ModelsabstractContext: predicting the number of defects in a defect backlog in a given time horizon can help allocate project resources and organize software development.Goal: to compare the accuracy of three defect backlog prediction methods in the context of large open-source (OSS) projects, i.e., ARIMA, Exponential Smoothing (ETS), and the state-of-the-art method developed at Ericsson AB (MS).Method: we perform a simulation study on a sample of 20 open-source projects to compare the prediction accuracy of the methods.Also, we use the Naïve prediction method as a baseline for sanity check.We use statistical inference tests and effect size coefficients to compare the prediction errors.Results: ARIMA, ETS, and MS were more accurate than the Naïve method.Also, the prediction errors were statistically lower for ETS than for MS (however, the effect size was negligible).Conclusions: ETS seems slightly more accurate than MS when predicting defect backlog size of OSS projects. Paulina Aniola, Sushant Kumar Pandey, Miroslaw Staron, Miroslaw Ochodek |
FedCSIS | 2 |
| 2023 | Design Patterns Understanding and Use in the Automotive Industry: An Interview Study
Sushant Kumar Pandey, Sivajeet Chand, Jennifer Horkoff, Miroslaw Staron |
PROFES (1) | 1 |
| 2022 | Comparing Input Prioritization Techniques for Testing Deep Learning AlgorithmsabstractDeep learning (DL) systems are becoming an essential part of software systems, so it is necessary to test them thoroughly. This is a challenging task since the test sets can grow over time as the new data is being acquired, and it becomes time-consuming. Input prioritization is necessary to reduce the testing time since prioritized test inputs are more likely to reveal the erroneous behavior of a DL system earlier during test execution. Input prioritization approaches have been rudimentary analyzed against each other, this study compares different input prioritization techniques regarding their effectiveness and efficiency. This work considers surprise adequacy, autoencoder-based, and similarity-based input prioritization approaches in the example of testing a DL image classification algorithms applied on MNIST, Fashion-MNIST, CIFAR-10, and STL-10 datasets. To measure effectiveness and efficiency, we use a modified APFD (Average Percentage of Fault Detected), and set up & execution time, respectively. We observe that the surprise adequacy is the most effective (0.785 to 0.914 APFD). The autoencoder-based and similarity-based techniques are less effective, with the performance from 0.532 to 0.744 APFD and 0.579 to 0.709 APFD, respectively. In contrast, the similarity-based and surprise adequacy-based approaches are the most and least efficient, respectively. The findings in this work demonstrate the trade-off between the considered input prioritization techniques to understanding their practical applicability for testing DL algorithms. Vasilii Mosin, Miroslaw Staron, Darko Durisic, Francisco Gomes de Oliveira Neto, Sushant Kumar Pandey, Ashok Chaitanya Koppisetty |
SEAA | 5 |
| 2021 | Machine learning based methods for software fault prediction: A survey
Sushant Kumar Pandey, Ravi Bhushan Mishra, Anil Kumar Tripathi |
Expert Syst. Appl. | 1 |
| 2021 | DNNAttention: A deep neural network and attention based architecture for cross project defect number prediction
Sushant Kumar Pandey, Anil Kumar Tripathi |
Knowl. Based Syst. | 1 |
| 2021 | An empirical study toward dealing with noise and class imbalance issues in software defect prediction
Sushant Kumar Pandey, Anil Kumar Tripathi |
Soft Comput. | 1 |
| 2020 | BPDET: An effective software bug prediction model using deep representation and ensemble learning techniques
Sushant Kumar Pandey, Ravi Bhushan Mishra, Anil Kumar Tripathi |
Expert Syst. Appl. | 1 |
| 2020 | Software defect prediction using K-PCA and various kernel-based extreme learning machine: an empirical studyabstractPredicting defects during software testing reduces an enormous amount of testing effort and help to deliver a high‐quality software system. Owing to the skewed distribution of public datasets, software defect prediction (SDP) suffers from the class imbalance problem, which leads to unsatisfactory results. Overfitting is also one of the biggest challenges for SDP. In this study, the authors performed an empirical study of these two problems and investigated their probable solution. They have conducted 4840 experiments over five different classifiers using eight NASA projects and 14 PROMISE repository datasets. They suggested and investigated the varying kernel function of an extreme learning machine (ELM) along with kernel principal component analysis (K‐PCA) and found better results compared with other classical SDP models. They used the synthetic minority oversampling technique as a sampling method to address class imbalance problems and k‐fold cross‐validation to avoid the overfitting problem. They found ELM‐based SDP has a high receiver operating characteristic curve over 11 out of 22 datasets. The proposed model has higher precision and F ‐score values over ten and nine, respectively, compared with other state‐of‐the‐art models. The Mathews correlation coefficient (MCC) of 17 datasets of the proposed model surpasses other classical models' MCC. Sushant Kumar Pandey, Deevashwer Rathee, Anil Kumar Tripathi |
IET Softw. | 1 |
| 2020 | BCV-Predictor: A bug count vector predictor of a successive version of the software system
Sushant Kumar Pandey, Anil Kumar Tripathi |
Knowl. Based Syst. | 1 |