EDBT 2026 Demo / reviewers in the wild / expert
Sandeep Kumar 0004
dblp:06/5441-4
· DBLP profile ↗
64ranked-venue papers
2as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 31 · 28 since 2021Artificial intelligence and machine learning · 13 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Client-Driven Greedy STOMP Connection Rebalancing Using Server-Reported Subscription Load for Adaptive Load Balancing
Manish Agrawal, Sandeep Kumar 0004, Subham Kumar, Ashok Shukla, Nandgopal Srinivasan |
ENASE (1) | 2 |
| 2026 | Towards Interpretable Ensemble Learning for Software Effort Estimation
Karishma Doshi, Suyash Shukla, Sandeep Kumar 0004 |
ENASE (1) | 3 |
| 2026 | An Encoder-Flexible Semantic Hypergraph Framework for API Recommendation in Mashup Development
Abhinav Jamwal, Dipender Singh, Manish Agrawal, Sandeep Kumar 0004 |
ENASE (1) | 4 |
| 2026 | An Extensive Empirical Investigation of Ensemble Learning for Software Defect Prediction
Megha Jayakumar, Suyash Shukla, Sandeep Kumar 0004 |
ENASE (1) | 3 |
| 2026 | Reinforcement Learning-Based Software Metric Selection for Defect Prediction
Dipender Singh, Abhinav Jamwal, Manish Agrawal, Sandeep Kumar 0004 |
ENASE (2) | 4 |
| 2026 | AMF-GR: Adaptive Matrix Factorization and Graph Fusion for Android Library Recommendation
Abhinav Jamwal, Sandeep Kumar 0004 |
SANER | 2 |
| 2026 | A Multi-Scale Hypergraph-Based Approach for Third-Party Library Recommendation in Mobile App DevelopmentabstractIn mobile app development, selecting the right third-party libraries (TPLs) is crucial to enhance functionality, improve code quality, and speed up the development process. However, recommending appropriate TPLs remains challenging due to the complexity of app-library interactions and the need to capture high-order relationships. Existing methods, such as collaborative filtering and graph-based approaches, often fail to adequately address these complexities. To address this challenge, we propose MsRec, a multiscale hypergraph neural network-based approach for TPL recommendation. MsRec uses hypergraphs to model interactions in groups of different sizes, leading to a detailed representation of the relationship between the application and the library. By modeling the strength, category, and functionality of interactions within each category, our framework improves the accuracy and diversity of recommendations. The multiscale hypergraph structure supports fine-grained relational reasoning, making it particularly effective in this context. Extensive experiments on real-world datasets demonstrate that MsRec outperforms current state-of-the-art methods, providing relevant and diverse TPL recommendations. Additionally, our model shows strong performance on various benchmarks, highlighting its ability to effectively handle complex app-library interactions. Abhinav Jamwal, Sandeep Kumar 0004 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2026 | FairGenerate: Enhancing Fairness through Synthetic Data Generation and Two-Fold Biased Labels RemovalabstractWe are increasingly using Machine Learning (ML) software to make autonomous decisions such as identifying credit risks, criminal sentencing, discharge of a patient, predicting heart diseases, hiring employees, and so on. However, the biased decision made by ML software against specific social groups based on protected/sensitive attributes (e.g., sex) has raised concern among Software Engineering (SE) and ML communities. ML software builds its decision logic from the training data. Consequently, if it is trained on biased data, it can make biased decisions. Previous studies have reported “biased labels” and “imbalanced data” as the root causes of biases in the training dataset. In this study, we propose FairGenerate, a pre-processing method that (a) balances the internal distribution of training datasets based on class labels and sensitive attributes by generating synthetic data samples using differential evolution and (b) identifies the biased labels through situation testing and removes them before and after synthetic data generation, hence the development of fair ML software. The experiments carried out in this study show that our proposed approach, FairGenerate, can attain markedly improved fairness (measured across various metrics) without compromising (original) the model performance over five benchmark methods. FairGenerate outperforms the fairness and performance tradeoff baseline set by the benchmarking tool Fairea in 65% of cases, compared to the state-of-the-art method, which achieves this in only 59% of cases. To promote open science, we provide all the scripts and data utilized in this work at https://github.com/xyz745/FairGenerate . Hem Chandra Joshi, Sandeep Kumar 0004 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2026 | An Approach for Cross-Project Defect Prediction Using Composite Distribution Alignment and Domain-Adversarial Neural Networks
Dipender Singh, Sandeep Kumar 0004 |
IEEE Trans. Reliab. | 2 |
| 2026 | Adaptive Spectral Clustering and Structural Alignment for Cross-Project Defect PredictionabstractCross-project defect prediction (CPDP) aims to identify defect-prone modules in a target project by leveraging data from external source projects. Over time, research has shifted from single-source to multi-source CPDP to increase data diversity and capture broader defect patterns. However, naively combining multiple sources can introduce substantial data divergence and degrade performance. A key limitation is that many methods overlook intra-project heterogeneity, treating each source project as uniform. Moreover, data selection typically relies on independent metric matching, which ignores structural relationships among software metrics. To address these limitations, we propose adaptive spectral clustering and structural alignment (ASCSA), a unified framework that combines adaptive spectral clustering and structural alignment. First, adaptive spectral clustering partitions each source project into coherent clusters, mitigating intraproject heterogeneity. Second, a structural alignment method selects source clusters that preserve higher-order metric relationships with the target, avoiding misleading matches based solely on marginal distributions. Finally, a deep neural network with maximum mean discrepancy (MMD) loss minimizes residual distribution gaps to enable effective knowledge transfer. In this study, we conduct comprehensive experiments on 25 widely used software projects. The results show that ASCSA consistently outperforms state-of-the-art methods, achieving improvements in AUC ranging from 5.26% to 23.08%, and in MCC from 15.62% to 68.18%. These findings highlight the effectiveness of jointly addressing intra-project heterogeneity, and structural alignment for reliable cross-project defect prediction. Dipender Singh, Sandeep Kumar 0004 |
IEEE Trans. Software Eng. | 2 |
| 2025 | Bridging AI and Human Knowledge: Towards a Deeper Understanding of Stack Overflow and ChatGPTabstractCommunity-driven forums like Stack Overflow (SO) have long established themselves as the go-to platform for developers seeking online help. Recently, ChatGPT, a powerful AI tool capable of generating high-level code and providing detailed explanations, has emerged as a strong alternative. While both platforms are valuable for developers, determining the best choice for specific use cases remains an open challenge. Although previous studies have examined the comparative merits of these platforms, the datasets used in such evaluations were limited. To bridge this gap, we introduce a four-dimensional benchmark dataset, ‘SEED’, that can facilitate a comprehensive analysis of ChatGPT and Stack Overflow. Our dataset comprises: (i) Developer Sentiments mined from 4161 comments from Reddit and SO meta-discussions, indicating community perceptions of both platforms, along with a manually labeled subset of 1,000 comments capturing developers’ expressed preferences; (ii) 3500 technical questions from SO, their accepted answers, and corresponding ChatGPT-generated responses for Efficacy (accuracy) benchmarking; (iii) An additional 200 deep learning-related SO posts, their accepted answers, and the corresponding ChatGPT answers to evaluate both these platforms on Energy efficiency parameters; (iv) 4,500 ChatGPT code snippets generated using tailor-made prompts designed to mimic SO answers for Detecting AI-code plagiarism. SEED can support diverse applications, including benchmarking AI-generated answers, evaluating energy efficiency in deep learning development, detecting AI plagiarism, and analyzing developer sentiment. By making this dataset publicly available, we lay the seed for advancing the research involving human-AI interaction in software engineering. Our dataset can be accessed at https://github.com/AnonymousResearch173/SEED. Aman Swaraj, Sandeep Kumar 0004 |
EASE | 2 |
| 2025 | Third-Party Library Recommendations Through Robust Similarity Measures
Abhinav Jamwal, Sandeep Kumar 0004 |
ENASE | 2 |
| 2025 | Towards an Approach for Project-Library Recommendation Based on Graph Normalization
Abhinav Jamwal, Sandeep Kumar 0004 |
ENASE | 2 |
| 2025 | VP-IAFSP: Vulnerability Prediction Using Information Augmented Few-Shot Prompting with Open Source LLMs
Mithilesh Pandey, Sandeep Kumar 0004 |
ENASE | 2 |
| 2025 | Detecting Adversarial Prompted AI-Generated Code on Stack Overflow: A Benchmark Dataset and an Enhanced Detection ApproachabstractAI-generated code has become an integral part of the mainstream developer workflow today. However, in community-driven platforms like Stack Overflow (SO), where trust, authorship, and credibility are important, it can lead to serious complications. While recent studies have focused on detecting AI-generated code, they have mostly worked with long code samples from repositories and assignments. In contrast, code snippets on SO are often small and context-specific, and thus may prove more challenging for detection. Moreover, another aspect overlooked in prior studies concerns recognizing adversarially prompted AI code deliberately crafted to resemble human-written code. To address these limitations, we have first introduced a large-scale dataset comprising 3500 pairs of SO and ChatGPT answers, along with a curated set of 4500 adversarially prompted AI responses. Next, we evaluate existing code language models over this newly curated dataset. Our evaluation shows that existing models perform well on standard AI answers but fail to detect adversarial ones. Finally, to improve detection, we propose an ensemble approach combining stylometric features of code along with the code embeddings. Our approach shows consistent improvements across multiple models and improves resistance to adversarial prompted code. Our overall findings open promising directions for future research into understanding the nuances of AI code detection with adversarial prompting and code stylometry. Aman Swaraj, Krishna Agarwal, Atharv Joshi, Sandeep Kumar 0004 |
ICSME | 4 |
| 2025 | StackPlagger: A System for Identifying AI-Code Plagiarism on Stack OverflowabstractIdentifying AI code plagiarism on technical forums like Stack Overflow (SO) is critical, as it can directly impact the platform’s trust and credibility. While previous studies have explored AI-generated code detection, they have focused on long, standalone samples from repositories and competitions. In contrast, SO snippets are often short, fragmented, and context-specific, which can make detection more challenging. Furthermore, existing methods have also not adequately addressed the concern of obfuscated or adversarially prompted code that are crafted to mimic human style and evade detection. To address these gaps, we first introduce a curated dataset of 8000 SO-ChatGPT snippet pairs generated using multiple adversarial prompts. While earlier methods solely relied on pre-trained models, we propose an ensemble approach combining stylometric features of code along with the pre-trained embeddings to improve detection performance. Finally, we deploy our fine-tuned model as a Google Chrome extension called ‘StackPlagger’, which can flag AI-generated code in SO answers and display AI confidence scores. Video demonstration and the associated artifacts of our tool can be found at https://youtu.be/6O9Urp2mvbI and https://github.com/harsh-g1/StackPlagger, respectively. Aman Swaraj, Harsh Goyal, Sumit Chadgal, Sandeep Kumar 0004 |
ASE | 4 |
| 2025 | DATSO: A Difficulty Assessment Tool for Stack Overflow QuestionsabstractStack Overflow is one of the prominent resources in the software community where developers seek technical assistance. Owing to its popularity, the platform witnesses questions of varying difficulty levels. Since not all practitioners are equally adept at addressing these queries, many queries remain unanswered or experience delays in receiving responses. However, the current framework of the site primarily categorizes the questions based on programming languages and specific topics, not considering the complexity level of the post. To bridge this gap, we introduce DATSO, a Difficulty Assessment Tool for Stack Overflow Questions. DATSO is a Chrome extension that assigns “difficulty level tags” to Stack Overflow posts by leveraging the textual and context-dependent features of the questions. Unlike previous works, which face issues like data imbalance and cold start problems, DATSO overcomes these limitations, achieving 72.3% accuracy on a benchmark dataset of Java questions, surpassing existing baseline works by 7%. The video demonstration and code repository can be found at https://youtu.be/VSu2Q0zb-X0 and https://github.com/krish10924/Difficulty-tag-Generator. Aman Swaraj, Neha Gujar, Manashree Kalode, Bhoomi Bonal, Krishna Agarwal, Sandeep Kumar 0004 |
SANER | 6 |
| 2025 | An approach to software defect prediction for small-sized datasets
Pravas Ranjan Bal, Suyash Shukla, Sandeep Kumar 0004 |
Appl. Intell. | 3 |
| 2025 | PredictPP: A Rank-Based Weighted Ensemble Model for Prediction of Software Project ProductivityabstractABSTRACT Software effort estimation (SEE) determines the effort necessary to develop software. The researchers have been tending to SEE issues since the 1960s, and several methods have been created until the formulation of the function point (FP) and constructive cost estimation (COCOMO) methods. However, these methods are only useful for procedurally developed software, not modern object‐oriented (OO) software. Because the use case is the widely used unit of an OO system, particularly in scenarios requiring structured and early‐stage effort estimation, using the use case point (UCP) approach will help get accurate results. The UCP approach consists of size estimation (in UCP) and effort estimation with calculated size. This study focuses on effort estimation when the size (in UCP) is already known. The productivity of a project is one of the main components for estimating effort from the given size. The classical SEE models based on UCP utilized a fixed number of productivity values. So, the validity of classical approaches is a subject of disapproval because of static productivity values. Purposefully, we proposed a rank‐based weighted ensemble model for productivity prediction that allows us to use flexible productivity values. We used learning techniques such as simple linear regression (SLR), Least Absolute Shrinkage and Selection Operator Regression (LR), ridge regression (RR), elastic net regression (ER), K‐nearest neighbor (KNN), decision tree (DT), support vector regression (SVR), multilayer perceptron (MLP), bagging, and adaptive boosting for productivity prediction and compared them with the proposed model. Further, we used existing UCP prediction models and compared the proposed approach with them. Suyash Shukla, Sandeep Kumar 0004 |
J. Softw. Evol. Process. | 2 |
| 2025 | An Approach for Cross Project Defect Prediction Using Identical Metrics Matching and Deep Neural NetworkabstractAdvancements in software defect prediction (SDP) to handle the scenario of no or limited historical data have introduced the concept of cross-project defect prediction (CPDP). CPDP using machine learning (ML) algorithms has been the staple research area for all software practitioners in the SDP domain. An important assumption in ML algorithms is that both train and test data must follow similar data distribution for better accuracy. These assumptions may hold in the within-project defect prediction (WPDP) scenario where both train and test data belong to the same project. However, it is impossible in the CPDP scenario where the train and test data belong to different projects. So, in the CPDP scenario, researchers tried to use a matched metrics approach to handling this issue. However, in this case, there may be an issue if only a small-sized source (train) dataset matches the data distribution with the target (test) dataset, leading to an insufficient training dataset. Hence, we have proposed a cross-project data preprocessing method, namely knowledge transfer from target data to source data using correlation (KTTSC), to handle this issue and hence to improve the CPDP accuracy of ML models. The experimental results demonstrate that using the dropout regularization-based deep neural network, k nearest neighbor, decision tree, logistic regression, and Naive Bayes classifiers with the proposed KTTSC method show an improvement of 22%, 17%, 23.2%, 13.5%, and 9.5%, respectively, in terms of average AUC scores as compared to the traditional CPDP method and an improvement in the range of 6.6% to 11.1% as compared to existing works on CPDP. Pravas Ranjan Bal, Sandeep Kumar 0004 |
IEEE Trans. Reliab. | 2 |
| 2024 | ME-SFP: A Mixture-of-Experts-Based Approach for Software Fault PredictionabstractIn the last two decades, many machine-learning-based works have been presented to build software fault prediction (SFP) models. The assessment of all these works showed that none of the machine learning classifiers could be generalized as the best-performing classifier in all prediction contexts. However, these techniques showed complementary behavior among them, which suggests combining their learning for improved performance. Various ensemble models explored in SFP have a major concern with the static weights assigned to the base learners for combining their decisions. In this article, we present a Mixture-of-Experts (MoE)-based approach named ME-SFP that uses the experts generated using a learning technique and a Gaussian mixture model as a gating function and a data partition technique. For the experimentation using the presented approach, we have shown the use of decision trees (DTs) as well as multilayer perceptrons (MLPs) as experts. We call them ME-SFP[DT] and ME-SFP[MLP]. We conduct a set of experiments on 35 publicly available software project datasets (PROMISE, JIRA, AEEEM, and Eclipse datasets) for SFP and measure the performance of built fault prediction models by using different measures such as F1-score, precision, recall, area under ROC curve (AUC), probability of false alarm, Mathews correlation coefficient, and G-means. Additionally, we perform Friedman's test and the Wilcoxon signed-rank sum test between the presented models, ensemble methods, baseline method, and individual learning techniques (DT and MLP). Results showed that ME-SFP[DT] and ME-SFP[MLP] produced improved results compared to individual techniques and ensemble methods. Aman Omer, Santosh Singh Rathore, Sandeep Kumar 0004 |
IEEE Trans. Reliab. | 3 |
| 2023 | Towards Automated Prediction of Software Bugs from Textual Description
Suyash Shukla, Sandeep Kumar 0004 |
ENASE | 2 |
| 2023 | Programming Language Identification in Stack Overflow Post Snippets with Regex Based Tf-Idf Vectorization over ANN
Aman Swaraj, Sandeep Kumar 0004 |
ENASE | 2 |
| 2023 | Self-Adaptive Ensemble-based Approach for Software Effort EstimationabstractSoftware Effort Estimation (SEE) is one of the most challenging tasks in software project management. In the literature, the researchers addressed the SEE problem in different ways, including models developed using machine learning (ML). Integrating individual models (Ensemble) is an active research topic in the ML domain, leading to improved performance than the individual models. Ensembling approaches such as bagging, boosting, and stacking have been extensively studied in the literature for improved SEE. The performance of an ensemble model largely depends upon tuning the hyperparameters of each learner and assigning the correct weight to each base learner. To this end, this work proposes a self-adaptive ensemble-based approach that combines hyperparameter tuning and weight assignment problems in a single step and solves them using optimization problems. The proposed model optimizes the hyperparameters of each learner using the particle swarm optimization (PSO) algorithm. Also, the weights of each base learner are optimized using Sequential Least SQuares Programming (SLSQP). An experimental analysis is conducted to evaluate the proposed model over different SEE datasets from PROMISE and ISBSG data repositories. The results show that the proposed ensemble approach performed better than the individual models. Further, we compared our proposed approach with other commonly used ensembling approaches in the literature. We found that the proposed ensemble approach provides better predictive performance than other ensembles. Suyash Shukla, Sandeep Kumar 0004 |
SANER | 2 |
| 2023 | An approach to occluded face recognition based on dynamic image-to-class warping using structural similarity index
Shadab Naseem, Santosh Singh Rathore, Sandeep Kumar 0004, Sugata Gangopadhyay, Ankita Jain |
Appl. Intell. | 3 |
| 2023 | Feature selection and clustering based web service selection using QoSs
Lalit Purohit, Santosh Singh Rathore, Sandeep Kumar 0004 |
Appl. Intell. | 3 |
| 2023 | Towards non-linear regression-based prediction of use case point (UCP) metric
Suyash Shukla, Sandeep Kumar 0004 |
Appl. Intell. | 2 |
| 2023 | Know-UCP: locally weighted linear regression based approach for UCP estimation
Suyash Shukla, Sandeep Kumar 0004 |
Appl. Intell. | 2 |
| 2023 | Towards ensemble-based use case point prediction
Suyash Shukla, Sandeep Kumar 0004 |
Softw. Qual. J. | 2 |
| 2023 | Concept Drift in Software Defect Prediction: A Method for Detecting and Handling the DriftabstractSoftware Defect Prediction (SDP) is crucial towards software quality assurance in software engineering. SDP analyzes the software metrics data for timely prediction of defect prone software modules. Prediction process is automated by constructing defect prediction classification models using machine learning techniques. These models are trained using metrics data from historical projects of similar types. Based on the learned experience, models are used to predict defect prone modules in currently tested software. These models perform well if the concept is stationary in a dynamic software development environment. But their performance degrades unexpectedly in the presence of change in concept (Concept Drift). Therefore, concept drift (CD) detection is an important activity for improving the overall accuracy of the prediction model. Previous studies on SDP have shown that CD may occur in software defect data and the used defect prediction model may require to be updated to deal with CD. This phenomenon of handling the CD is known as CD adaptation. It is observed that still efforts need to be done in this direction in the SDP domain. In this article, we have proposed a pair of paired learners (PoPL) approach for handling CD in SDP. We combined the drift detection capabilities of two independent paired learners and used the paired learner (PL) with the best performance in recent time for next prediction. We experimented on various publicly available software defect datasets garnered from public data repositories. Experimentation results showed that our proposed approach performed better than the existing similar works and the base PL model based on various performance measures. Arvind Kumar Gangwar, Sandeep Kumar 0004 |
ACM Trans. Internet Techn. | 2 |
| 2023 | A QoS-Aware Clustering Based Multi-Layer Model for Web Service SelectionabstractThe rapid proliferation of new web services over the last decade has led to an increase in functionally identical services, making the service selection system more challenging. In this work, we propose a web service selection model consisting of two layers, which we call CPSky. The upper layer, called the Prefilter layer, filters and allows only potent services to participate in the selection process. Pruning and clustering based on the quality of service form the basis of prefiltering. The bottom layer, called the Selection layer, uses the proposed Skyline-Plus approach to select the appropriate web service. The existing skyline technique always generates the same set of skyline services and does not consider the end-user requested QoS. Additionally, the skyline results in a non-dominated set of services without any ordering of services. To address these issues, we propose a modified skyline, which we call Skyline-Plus. The Selection layer also identifies replaceable web services using the Pearson correlation coefficient. The proposed approach's efficacy is validated through experimental evaluation on a real-world dataset using four performance evaluation parameters. The experimental results show that the proposed approach performs better than existing similar approaches in terms of efficiency and end-user satisfaction. Lalit Purohit, Santosh Singh Rathore, Sandeep Kumar 0004 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | A Data Transfer and Relevant Metrics Matching Based Approach for Heterogeneous Defect PredictionabstractHeterogeneous defect prediction (HDP) is a promising research area in the software defect prediction domain to handle the unavailability of the past homogeneous data. In HDP, the prediction is performed using source dataset in which the independent features (metrics) are entirely different than the independent features of target dataset. One important assumption in machine learning is that independent features of the source and target datasets should be relevant to each other for better prediction accuracy. However, these assumptions do not generally hold in HDP. Further in HDP, the selected source dataset for a given target dataset may be of small size causing insufficient training. To resolve these issues, we have proposed a novel heterogeneous data preprocessing method, namely, Transfer of Data from Target dataset to Source dataset selected using Relevance score (TDTSR), for heterogeneous defect prediction. In the proposed approach, we have used chi-square test to select the relevant metrics between source and target datasets and have performed experiments using proposed approach with various machine learning algorithms. Our proposed method shows an improvement of at least 14% in terms of AUC score in the HDP scenario compared to the existing state of the art models. Pravas Ranjan Bal, Sandeep Kumar 0004 |
IEEE Trans. Software Eng. | 2 |
| 2022 | A Methodology for Detecting Programming Languages in Stack Overflow Questions
Aman Swaraj, Sandeep Kumar 0004 |
ICSOFT | 2 |
| 2022 | Predicting reliability of software in industrial systems using a Petri net based approach: A case study on a safety system used in nuclear power plant
Sumit, Sandeep Kumar 0004, Lalit Kumar Singh, Alok Mishra 0001 |
Inf. Softw. Technol. | 3 |
| 2022 | Applications of deep learning for phishing detection: a systematic literature review
Cagatay Catal, Görkem Giray, Bedir Tekinerdogan, Sandeep Kumar 0004, Suyash Shukla |
Knowl. Inf. Syst. | 4 |
| 2022 | Towards a super-resolution based approach for improved face recognition in low resolution environment
Nalin Singh, Santosh Singh Rathore, Sandeep Kumar 0004 |
Multim. Tools Appl. | 3 |
| 2021 | An Extreme Learning Machine based Approach for Software Effort Estimation
Suyash Shukla, Sandeep Kumar 0004 |
ENASE | 2 |
| 2021 | A Stacking Ensemble-based Approach for Software Effort Estimation
Suyash Shukla, Sandeep Kumar 0004 |
ENASE | 2 |
| 2021 | An empirical study of ensemble techniques for software fault prediction
Santosh Singh Rathore, Sandeep Kumar 0004 |
Appl. Intell. | 2 |
| 2021 | Software fault prediction based on the dynamic selection of learning technique: findings from the eclipse project study
Santosh Singh Rathore, Sandeep Kumar 0004 |
Appl. Intell. | 2 |
| 2021 | A Classification Based Web Service Selection ApproachabstractSelection of an appropriate web service fulfilling the requirements of the end user is a challenging task. Most of the existing systems use Quality of Service (QoS) as predominant parameter for web service selection, without any preprocessing or filtering. These systems consider all of the candidate web services during selection process and require unnecessary processing of those web services which are far below the expectations of the end user. In this work, an approach for web service selection based on QoS parameters is proposed. The proposed method starts with prefiltering of candidate web services using classification technique. An improved PROMETHEE method, we call it as PROMETHEE Plus, is applied to most eligible web services and Maximizing Deviation Method based hybrid weight evaluation mechanism is adopted. Top-k web services matching closely with the QoS requirements of the end user are selected. Experiments on the dataset of real world web services are conducted. Experimental results show that our approach performs better in terms of end user satisfaction and efficiency with reference to the existing similar approaches. Lalit Purohit, Sandeep Kumar 0004 |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | WR-ELM: Weighted Regularization Extreme Learning Machine for Imbalance Learning in Software Fault PredictionabstractImbalanced data is a significant issue in software fault prediction. It is very challenging for software engineers to handle imbalanced software fault data for the early prediction of software faults. In the last two decades, many researchers have used synthetic minority oversampling technique (SMOTE), SMOTE for regression and other such techniques to preprocess the imbalanced software fault data. However, these preprocessing techniques do not produce consistently good accuracy, especially in inter release, and cross project fault prediction. The learning of imbalanced fault data for prediction of the number of software faults has not been explored in depth so far. To deal with this scenario, we have explored an efficient machine learning technique, namely extreme learning machine (ELM) for prediction of the number of software faults. Furthermore, a new variant of ELM, namely weighted regularization ELM, is proposed to generalize the imbalanced data to balanced data. To validate the proposed imbalanced learning model, we have used 26 open source PROMISE software fault datasets and three prediction scenarios, intra release, inter release, and cross project. We have conducted the experiments for prediction of the number of faults. The experimental results showed that the proposed approach led to improved performance. Pravas Ranjan Bal, Sandeep Kumar 0004 |
IEEE Trans. Reliab. | 2 |
| 2019 | Replaceability Based Web Service Selection ApproachabstractWeb services are important components of modern day software. Many services are composed on the fly to offer the required business functionality. Web services run in very dynamic environment and are prone to failure or run time change in service quality. On the occurrence of failure or unavailability of any component service, replacement by an equivalent service is required. In this paper, a web service selection approach by taking into cognizance the replaceability of service as the basis of selection is used. At the time of selection, the replaceability of a web service is determined using improved PROMETHEE method. The task of obtaining replacement service at run time is very expensive and challenging. The experimental results depict the improvements achieved in the results of selection. The experiments are conducted by taking QoS dataset of real world web services. Lalit Purohit, Sandeep Kumar 0004 |
HiPC | 2 |
| 2019 | Clustering Based Approach for Web Service Selection Using Skyline ComputationsabstractWeb services are useful to automate a task. Along with automation of task, efficiency improvement is another important challenge for researchers of web service community. To improve the overall execution efficiency of web service based system, the input to selection process needs to be preprocessed. In this work, the clustering is applied to candidate web services to determine similar services on the basis of QoS information. A systematic analysis is done to evaluate the performance of three clustering techniques using Dunn index and average distance measure. The best performing clustering technique is applied on candidate web services. The most prominent set of web services is considered for skyline based selection. To perform various experiments, a QoS dataset based on real world web services is used. It is evident from the results of experimentation that the proposed approach is better than existing similar approaches for web service selection. Lalit Purohit, Sandeep Kumar 0004 |
ICWS | 2 |
| 2019 | Development of Machine Learning Based Approach for Computing Optimal Vegetation Index with The Use of Sentinel-2 And Drone DataabstractSatellite imagery has been used in most of the applications to estimate, monitor and predict various situations and hurdles in the current scenario. According to the application, the satellite data is chosen on the basis of its spatial, spectral and temporal characteristics. From agriculture to the road highways construction and their health monitoring, it is playing a major role by providing its features details from a high level to medium and low level. In the context of agriculture, there are many ways to find crop health like visual interpretation, health parameter, existing vegetation indices, etc. For the purpose of monitoring and predicting, we require band values from the satellite data which may lack in providing accurate values because of the several parameters that affect it by which there is a possibility of error because they have to be calibrated. Therefore in this paper, a machine learning based calibration approach is proposed by considering multispectral drone data as in-situ data. The methodology is developed and tested for Sentinel-2 data. It is observed that the proposed calibration technique has good potential to calibrate satellite data. Ankush Agarwal, Sandeep Kumar 0004, Dharmendra Singh |
IGARSS | 2 |
| 2019 | Applicability of Neural Network Based Models for Software Effort EstimationabstractEffort Estimation is a very challenging task in the software development life cycle. Inaccurate estimations may cause the client dissatisfaction and thereby, decrease the quality of the product. Considering the problem of software cost and effort prediction, it is conceivable to call attention to that the estimation procedure considers the qualities present in the data set, as well as the aspects of the environment in which the model is embedded. Existing literatures have the instances where machine learning techniques such as Linear Regression (LR), Support Vector Machine (SVM), K-Nearest Neighbor (KNN) have been used to estimate the effort required to develop any software. Yet it is quite uncertain for any particular model to perform well with all the data sets. Most of the research is based on the dataset of any single organization. Consequently, the results obtained through these models cannot be generalized. So, the main objectives of this research are: i) to use different data preparation techniques such as selection, cleaning, and transformation to improve the quality of data set given to the model ii) to use other machine learning models such as Multi-Layer Perceptron Neural Network (MLPNN), Probabilistic Neural Network (PNN), and Recurrent Neural Network (RNN) to increase the performance of software effort estimation process iii) to use different optimization techniques to tune the parameters of machine learning models iv) to use ensemble methods to improve the accuracy of software effort estimation process. In this study, first, we found out the most influential attributes in the Desharnais data set, then, MLPNN has been applied on reduced data set with to improve the accuracy of software effort estimation. Then, the performance of the MLPNN model is compared with LR, SVM and KNN models in the literature to find the best model fitting this dataset. Results obtained from the study demonstrate that some of the variables are more important in comparison to others for effort estimation. Also among the various models used in this study, the best-obtained R2 value is 79 % for the MLPNN model. Suyash Shukla, Sandeep Kumar 0004 |
SERVICES | 2 |
| 2019 | Analyzing Effect of Ensemble Models on Multi-Layer Perceptron Network for Software Effort EstimationabstractEffort Estimation is a very challenging task in the software development life cycle. Inaccurate estimations may cause client dissatisfaction and thereby, decrease the quality of the product. Considering the problem of software cost and effort estimation, it is conceivable to call attention to that the estimation procedure considers the qualities present in the data set, as well as the aspects of the environment in which the model is embedded. Existing literature have the instances where machine learning techniques have been used to estimate the effort required to develop any software. Yet it is quite uncertain for any particular model to perform well with all the data sets. In this paper, Multi-Layer Perceptron (MLPNN) and its ensembles are explored in order to improve the performance of software effort estimation process. Firstly, MLPNN, Ridge-MLPNN, Lasso-MLPNN, Bagging-MLPNN, and AdaBoost-MLPNN models are developed and, then, the performance of these models are compared on the basis of R2score to find the best model fitting this dataset. Results obtained from the study demonstrate that the R2score of AdaBoost-MLPNN is 82.213%, which is highest among all the models. Suyash Shukla, Sandeep Kumar 0004, Pravas Ranjan Bal |
SERVICES | 2 |
| 2019 | An Approach for the Prediction of Number of Software Faults Based on the Dynamic Selection of Learning TechniquesabstractDetermining the most appropriate learning technique(s) is vital for the accurate and effective software fault prediction (SFP). Earlier techniques used for SFP have reported varying performance for different software projects and none of them has always reported the best performance across different projects. The problem of varying performance can be solved by using an approach, which partitions the fault dataset into different module subsets, trains learning techniques for each subset, and integrates the outcomes of all the learning techniques. This paper presents an approach that dynamically selects learning techniques to predict the number of software faults. For a given testing module, the presented approach first locates its neighbor module subset that contained modules similar to testing module using a distance function and then chooses the best learning technique in the region of that module subset to make the prediction for testing module. The learning technique is selected based on its past performance in the region of module subset. We have performed an evaluation of the proposed approach using fault datasets garnered from the PROMISE data repository and Eclipse bug data repository. Experimental results showed that the proposed approach led to an improved performance when predicting the number of faults in software systems. Santosh Singh Rathore, Sandeep Kumar 0004 |
IEEE Trans. Reliab. | 2 |
| 2018 | Extreme Learning Machine based Linear Homogeneous Ensemble for Software Fault Prediction
Pravas Ranjan Bal, Sandeep Kumar 0004 |
ICSOFT | 2 |
| 2018 | Cross Project Software Defect Prediction using Extreme Learning Machine: An Ensemble based Study
Pravas Ranjan Bal, Sandeep Kumar 0004 |
ICSOFT | 2 |
| 2018 | SWAQ: a Semantic Web Application Quality Evaluation FrameworkabstractSemantic Web Applications (SWAs) differ from Web applications and software as they capture semantics and support heterogeneous information. Therefore, evaluation of SWA’s quality should be treated differently. Quality framework for SWAs should not only retain the attributes of Web applications and software that are relevant to SWAs, but also accommodate attributes unique to SWAs. We propose SWAQ framework comprising metrics for evaluation of SWA quality attributes, which have been validated using standard criteria for quality models. SWAQ uses analytic hierarchy process for multiple criteria decision making to rank SWAs. Moreover, a fuzzy system is employed for handling uncertainty involved in quality evaluation. Niyati Baliyan, Sandeep Kumar 0004 |
J. Exp. Theor. Artif. Intell. | 2 |
| 2018 | FriendShare: A secure and reliable framework for file sharing on network
Sandeep Kumar 0004 |
J. Netw. Comput. Appl. | 2 |
| 2017 | Towards an ensemble based system for predicting the number of software faults
Santosh Singh Rathore, Sandeep Kumar 0004 |
Expert Syst. Appl. | 2 |
| 2017 | Ontology Cohesion and Coupling MetricsabstractCohesion and coupling metrics for a modular ontology provide an insight into the efficiency of the modularization technique employed. Most of such metrics are either syntax based or do not account for ontology hierarchy. The authors propose cohesion and coupling metrics with continuous scale of measurement for modular ontology via deducing the degree of dependence among ontology components, thereby, handling structural aspects of ontology. The proposed metrics further handle the subtle differences between the type of links. The metrics have been analytically validated using well-established frameworks and the experimental results compared with the existing works, in terms of their cohesion and coupling values as well as the features of ontologies that they can measure. The proposed metrics are unambiguous and can be applied to any type of ontology in order to facilitate the assessment of the quality of an ontology or a partition thereof. Sandeep Kumar 0004, Niyati Baliyan, Shriya Sukalikar |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2017 | Research-oriented teaching of PDC topics in integration with other undergraduate courses at multiple levels: A multi-year report
Sandeep Kumar 0004 |
J. Parallel Distributed Comput. | 1 |
| 2017 | Linear and non-linear heterogeneous ensemble methods to predict the number of faults in software systems
Santosh Singh Rathore, Sandeep Kumar 0004 |
Knowl. Based Syst. | 2 |
| 2017 | Predicting information diffusion probabilities in social networks: A Bayesian networks based approach
Devesh Varshney, Sandeep Kumar 0004, Vineet Gupta 0002 |
Knowl. Based Syst. | 2 |
| 2017 | An Approach to Classify Tall Vegetation and Urban Using Deoriented PALSAR ImageabstractResearchers are attempting to classify the fully polarimetric synthetic aperture radar data utilizing diverse methods based on polarimetric indices, decomposition, and image analysis. Although, all these techniques have great potential, and able to segregate distinct land cover classes, but still there occurs ambiguity in classifying urban and tall vegetation classes as they both show similar kind of double-bounce scattering characteristics. In the past, Yamaguchi has modified the polarization process by applying the deorientation effect, which has enhanced the double-bounce characteristic of similar classes like urban and tall vegetation. Based on this, an attempt has been made in this letter to remove the uncertainty between the two classes by proposing a deorientation feature-based classification algorithm which could segregate urban and tall vegetation in an unsupervised way. In this, along with polarimetric, other features, namely, color, texture, and wavelets, have been critically analyzed on the deoriented image as each feature type has its own points of interest and hindrances. The result obtained from the proposed technique has shown good accuracy rate for urban and tall vegetation classification. Dharmendra Singh, Sandeep Kumar 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Towards an ontology based framework for searching multimedia contents on the web
Shikhar Shrivastav, Sandeep Kumar 0004 |
Multim. Tools Appl. | 2 |
| 2017 | An empirical study of some software fault prediction techniques for the number of faults prediction
Santosh Singh Rathore, Sandeep Kumar 0004 |
Soft Comput. | 2 |
| 2016 | Class wise optimal feature selection for land cover classification using SAR dataabstractConventional methods for classifying SAR data, such as H-α decomposition, Wishart classifier etc. are quite complex and classifies data only on the basis of polarimetric information. With the advent of distinct feature types, their role in land cover classification using SAR data could be analysed. For the sake of classification, researchers are extracting and combining several features in order to obtain the best attainable accuracy. But the usage of several feature type is not only increasing the computational complexity, but also the salience of each of the feature type remains unhighlighted. Hence, it became difficult to analyse that which feature type are best suitable for classification and selection of suitable features for land cover classification is challenging as each feature has its own significance level. Therefore, in this paper class wise, optimal feature selection for land cover classification has been performed using SAR data. For optimal feature selection, four types of feature set polarimetric features, texture features, color features and wavelet features have been examined. For class wise feature subset selection separability index criteria and classification results obtained using Naive Bayes classifier has been utilized. With the proposed methodology overall 10 features has been selected among the total 37 feature analysed with fine land cover classification accuracy of 91%. Sandeep Kumar 0004, Akanksha Garg, Dharmendra Singh, N. S. Rajput 0001 |
IGARSS | 2 |
| 2015 | Summarizing Customer Reviews through Aspects and Contexts
Prakhar Gupta, Sandeep Kumar 0004, Kokil Jaidka |
CICLing (2) | 2 |
| 2015 | An efficient use of random forest technique for SAR data classificationabstractIn the past SAR data has been proven as a great source for land cover characterization. For classification purpose many individual methods has been used, but single method are likely to undergo high variance or biasness depending on the base used for classification. Hence, in this paper random forest classification technique has been used for SAR data classification into different land cover classes (urban, water, vegetation and bare soil) which minimizes the diversity amongst the fragile classifiers and produce more accurate predictions. In this regard, an attempt has been made to fuse, four types of measures, namely texture features, SAR observable, statistical features and color features using random forest classifier for land cover classification. The results show that the resultant classified image has better accuracy in comparison to the individual method. Dharmendra Singh, Keshava P. Singh, Sandeep Kumar 0004 |
IGARSS | 4 |
| 2014 | Modeling Information Diffusion in Social Networks Using Latent Topic Information
Devesh Varshney, Sandeep Kumar 0004, Vineet Gupta 0002 |
ICIC (1) | 2 |