VLDB 2026 Research / reviewers in the wild / expert
Santosh Singh Rathore
dblp:119/1678
· DBLP profile ↗
29ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0003-2087-1666ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Human Values Perspective on Playability Issues of Mobile GamesabstractThe rise of mobile devices and gaming platforms has transformed the mobile gaming industry. A game’s playability, driven by functionality, usability, and satisfaction, directly affects player experience. Online app stores provide reviews that reveal gameplay strengths, issues, and user concerns, enabling developers to assess popularity and address problems. Yet, game development relies heavily on internal play-testing, often overlooking real user experiences. This study analyses mobile game playability through the lens of human values neglected in user reviews. Using Schwartz’s theory of human values and Sánchez’s playability model, we examined 20,346 reviews from the top 15 Google Play Store games. A fine-grained analysis identified 42 functionalities, concerns, and negligence issues. Results show nearly 30% of reviews report violations of human values, significantly affecting playability. Among value categories, Socialism emerged as the most neglected, while Emotion was the least. These findings highlight critical gaps in current mobile game design practices. Swapnil Thakar, Saurabh Tiwari 0001, Santosh Singh Rathore |
Int. J. Hum. Comput. Interact. | 3 |
| 2026 | Systematic literature review on software code smell detection approaches
Praveen Singh Thakur, Satyendra Singh Chouhan, Santosh Singh Rathore, Jitendra Parmar |
J. Syst. Softw. | 3 |
| 2026 | COSTAR: Software Code Smell Detection Through Tree-Based Abstract RepresentationabstractCode smells are suboptimal code structures that increase software maintenance costs and are challenging to detect manually. Researchers have explored automatic code smell detection using Machine Learning (ML) methods, which rely heavily on static code metrics or source code representation. Static code metrics often rely on structural attributes such as lines of code, cyclomatic complexity, or comment density. However, these metrics do not always reflect true code complexity and provide only quantitative insights without inherently detecting poor coding practices. In contrast, representations like Abstract Syntax Trees (ASTs) focus on the structural and syntactic elements of code, capturing hierarchical and contextual relationships within the source code. This enables precise identification of code structures such as loops, function calls, and conditionals, which are essential for detecting code smells. This paper introduces COSTAR (Code Smell Detection through Tree-based Abstract Representation), a source code representation technique using Abstract Syntax Trees (AST) to uniquely represent each source code instance. COSTAR captures the hierarchical structure of the source code by extracting all paths from the root to individual nodes within the AST. By employing a pretrained Sentence-BERT (SBERT) embedding model, COSTAR generates vectors for each extracted path. The subsequent calculation of the mean of these vectors yields a precise and comprehensive source code representation. Extensive experiments were conducted to validate COSTAR's performance using various ML techniques on four benchmark MLCQ code smell datasets: Data Class, God Class (Blob), Feature Envy, and Long Method. Various performance metrics have been employed to evaluate the model's performance. The experimental results indicate that COSTAR enhances the performance of the code smell detection model compared to existing methods. An improvement in the f1-score ranging from 0.03 (Long Method) to 0.19 (Feature Envy) was observed. Furthermore, a comparison of COSTAR with state-of-the-art methods demonstrated that it outperformed approaches like Code2Vec and CuBERT in code smell detection. Praveen Singh Thakur, Mahipal Jadeja, Satyendra Singh Chouhan, Santosh Singh Rathore |
IEEE Trans. Reliab. | 4 |
| 2025 | Leveraging LLMs for Requirements Engineering Education: How to Approach?abstractRequirements Engineering (RE) is a key yet often challenging phase that demands a good understanding of stakeholder needs, domain, elicitation methods, and documentation practices. Teaching RE is challenging due to the complexity of technical processes paired with critical soft skills. Role-Playing is a common and efficient technique in RE Education (REE), strengthening students’ comprehension of stakeholder interaction and requirement elicitation. The quality and consistency of traditional role-playing are nevertheless susceptible to the instructor’s facilitation abilities, students’ role-playing capabilities, and the dynamic of each group. The emergence of Generative AI (GenAI) and subsequent Large Language Models (LLMs) has introduced new opportunities for providing tailored support to students beyond conventional learning resources. In this paper, we explore the potential of LLMs for REE and how LLMs can assist in teaching RE concepts to students. We have conducted a pilot study with forty-six students to explore teaching RE concepts and developing pedagogy in REE by assigning the role of co-analyst to the LLM. Our results show that LLMs help students understand problems from various perspectives, providing a realistic view of the underlying complexities and alternative solutions for the RE tasks. Saurabh Tiwari 0001, Santosh Singh Rathore |
RE | 2 |
| 2025 | ATE-FS: An Average Treatment Effect-Based Feature Selection Technique for Software Fault PredictionabstractIn software development, software fault prediction (SFP) models aim to identify code sections with a high likelihood of faults before the testing process. SFP models achieve this by analyzing data about the structural properties of the software’s previous versions. Consequently, the accuracy and interpretation of SFP models depend heavily on the chosen software metrics and how well they correlate with patterns of fault occurrence. Previous research has explored improving SFP model performance through feature selection (metric selection). Yet inconsistencies in conclusions arose due to the presence of inconsistent and correlated software metrics. Relying solely on correlations between metrics and faults makes it difficult for developers to take actionable steps, as the causal relationships remain unclear. To address this challenge, this work investigates the use of Causal Inference (CI) methods to understand the causal relationships between software project characteristics, development practices, and the fault-proneness of code sections. We propose a CI-based technique called Average Treatment Effect for Feature Selection (ATE-FS). This technique leverages the causal inference concept to quantify the cause-and-effect relationships between software metrics and fault-proneness. ATE-FS utilizes Average Treatment Effect (ATE) features to identify code metrics that are most suitable for building SFP models. These ATE features capture the causal impact of a metric on fault-proneness. Through an experimental analysis involving twenty-seven SFP datasets, we validate the performance of ATE-FS. We further compare its performance with other state-of-the-art feature selection techniques. The results demonstrate that ATE-FS achieves a significant performance for fault prediction. Additionally, ATE-FS improved consistency in feature selection across diverse SFP datasets. Akshat Mangal, Santosh Singh Rathore |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2024 | Beyond Text: Multimodal Credibility Assessment Approaches for Online User-Generated ContentabstractUser-generated content (UGC) is increasingly becoming prevalent on various digital platforms. The content generated on social media, review forums, and question–answer platforms impacts a larger audience and influences their political, social, and other cognitive abilities. Traditional credibility assessment mechanisms involve assessing the credibility of the source and the text. However, with the increase in how user content can be generated and shared (audio, video, and images), multimodal representation of UGC has become increasingly popular. This article reviews the credibility assessment of UGC in various domains, particularly identifying fake news, suspicious profiles, and fake reviews and testimonials, focusing on both textual content and the source of the content creator. Next, the concept of multimodal credibility assessment is presented, which also includes audio, video, and images in addition to text. After that, the article presents a systematic review and comprehensive analysis of work done in the credibility assessment of UGC considering multimodal features. Additionally, the article provides extensive details on the publicly available multimodal datasets for the credibility assessment of UGC. In the end, the research gaps, challenges, and future directions in assessing the credibility of multimodal UGC are presented. Monika Choudhary, Satyendra Singh Chouhan, Santosh Singh Rathore |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | ME-SFP: A Mixture-of-Experts-Based Approach for Software Fault PredictionabstractIn the last two decades, many machine-learning-based works have been presented to build software fault prediction (SFP) models. The assessment of all these works showed that none of the machine learning classifiers could be generalized as the best-performing classifier in all prediction contexts. However, these techniques showed complementary behavior among them, which suggests combining their learning for improved performance. Various ensemble models explored in SFP have a major concern with the static weights assigned to the base learners for combining their decisions. In this article, we present a Mixture-of-Experts (MoE)-based approach named ME-SFP that uses the experts generated using a learning technique and a Gaussian mixture model as a gating function and a data partition technique. For the experimentation using the presented approach, we have shown the use of decision trees (DTs) as well as multilayer perceptrons (MLPs) as experts. We call them ME-SFP[DT] and ME-SFP[MLP]. We conduct a set of experiments on 35 publicly available software project datasets (PROMISE, JIRA, AEEEM, and Eclipse datasets) for SFP and measure the performance of built fault prediction models by using different measures such as F1-score, precision, recall, area under ROC curve (AUC), probability of false alarm, Mathews correlation coefficient, and G-means. Additionally, we perform Friedman's test and the Wilcoxon signed-rank sum test between the presented models, ensemble methods, baseline method, and individual learning techniques (DT and MLP). Results showed that ME-SFP[DT] and ME-SFP[MLP] produced improved results compared to individual techniques and ensemble methods. Aman Omer, Santosh Singh Rathore, Sandeep Kumar 0004 |
IEEE Trans. Reliab. | 2 |
| 2023 | An approach to occluded face recognition based on dynamic image-to-class warping using structural similarity index
Shadab Naseem, Santosh Singh Rathore, Sandeep Kumar 0004, Sugata Gangopadhyay, Ankita Jain |
Appl. Intell. | 2 |
| 2023 | Feature selection and clustering based web service selection using QoSs
Lalit Purohit, Santosh Singh Rathore, Sandeep Kumar 0004 |
Appl. Intell. | 2 |
| 2023 | TextConvoNet: a convolutional neural network based architecture for text classification
Sanskar Soni, Satyendra Singh Chouhan, Santosh Singh Rathore |
Appl. Intell. | 3 |
| 2023 | An attention-based deep learning model for credibility assessment of online health informationabstractAbstract With the surge of searching and reading online health‐based articles, maintaining the quality and credibility of online health‐based articles has become crucial. The circulation of deceptive health information on numerous social media sites can mislead people and can potentially cause adverse effects on people's health. To address these problems, this work uses deep learning approaches to automate the assessment and scoring of online health‐related articles' credibility. The paper proposed an Attention‐based Recurrent Multichannel Convolutional Neural Network (ARMCNN) model. The proposed model incorporates a BiLSTM layer, a multichannel CNN layer, and an attention layer and predicts the credibility of online health information. To perform a reliable evaluation of the presented model, we utilize the health articles reviewed by the experts, compiled in a labeled dataset termed “Pubhealth,” which consists of thousands of health articles. The results are evaluated using five performance measures, accuracy, precision, recall, f1‐score, and area under the ROC curve (AUC). Furthermore, we extensively compared the proposed model with different deep learning and machine learning models such as Long short‐term memory (LSTM), Bidirectional LSTM, CNN (Convolutional neural network), and RNN‐CNN. The experimental results showed that the proposed model produced state‐of‐the‐art performance on the used dataset by achieving an accuracy of 0.88, precision of 0.92, recall of 0.87, f1‐score of 0.90, and AUC of 0.94. Further, the proposed model yielded better performance than other benchmarked techniques for the credibility assessment of online health articles. Swarup Padhy, Santosh Singh Rathore |
Comput. Intell. | 2 |
| 2023 | Implicit and explicit mixture of experts models for software defect prediction
Aditya Shankar Mishra, Santosh Singh Rathore |
Softw. Qual. J. | 2 |
| 2023 | A QoS-Aware Clustering Based Multi-Layer Model for Web Service SelectionabstractThe rapid proliferation of new web services over the last decade has led to an increase in functionally identical services, making the service selection system more challenging. In this work, we propose a web service selection model consisting of two layers, which we call CPSky. The upper layer, called the Prefilter layer, filters and allows only potent services to participate in the selection process. Pruning and clustering based on the quality of service form the basis of prefiltering. The bottom layer, called the Selection layer, uses the proposed Skyline-Plus approach to select the appropriate web service. The existing skyline technique always generates the same set of skyline services and does not consider the end-user requested QoS. Additionally, the skyline results in a non-dominated set of services without any ordering of services. To address these issues, we propose a modified skyline, which we call Skyline-Plus. The Selection layer also identifies replaceable web services using the Pearson correlation coefficient. The proposed approach's efficacy is validated through experimental evaluation on a real-world dataset using four performance evaluation parameters. The experimental results show that the proposed approach performs better than existing similar approaches in terms of efficiency and end-user satisfaction. Lalit Purohit, Santosh Singh Rathore, Sandeep Kumar 0004 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Securing multimedia videos using space-filling curves
Debanjan Sadhya, Santosh Singh Rathore, Amitesh Singh Rajput |
Multim. Tools Appl. | 2 |
| 2022 | Towards a super-resolution based approach for improved face recognition in low resolution environment
Nalin Singh, Santosh Singh Rathore, Sandeep Kumar 0004 |
Multim. Tools Appl. | 2 |
| 2022 | Generative Oversampling Methods for Handling Imbalanced Data in Software Fault PredictionabstractImbalanced software fault datasets, having fewer faulty modules than the nonfaulty modules, make accurate fault prediction difficult. It is challenging for software practitioners to handle imbalanced fault data during software fault prediction (SFP). Earlier, several researchers have applied oversampling techniques such as synthetic minority oversampling techniques and others for imbalanced learning in SFP. However, most of these techniques resulted in overfitted prediction models. This article presents generative oversampling methods to handle imbalanced data problems in the SFP. Using the generative adversarial network (GAN) based approach, the presented methods generate synthetic samples of the faulty modules to balance the proportion of faulty and nonfaulty modules in the fault datasets. Further, SFP models are built on the processed fault datasets using different machine learning techniques. Experimental validation of the presented oversampling methods is done on 18 fault datasets gathered from PROMISE, JIRA, Eclipse data repositories, and precision, recall, f1-score, and AUC are used as evaluation measures. We extensively compared presented oversampling methods with various state-of-the-art class imbalance techniques and baseline models. The experimental results evidenced that the presented methods improved fault prediction performance and yielded better performance than the state-of-the-art class imbalance techniques. Santosh Singh Rathore, Satyendra Singh Chouhan, Dixit Kumar Jain, Aakash Gopal Vachhani |
IEEE Trans. Reliab. | 1 |
| 2021 | An empirical study of ensemble techniques for software fault prediction
Santosh Singh Rathore, Sandeep Kumar 0004 |
Appl. Intell. | 1 |
| 2021 | Software fault prediction based on the dynamic selection of learning technique: findings from the eclipse project study
Santosh Singh Rathore, Sandeep Kumar 0004 |
Appl. Intell. | 1 |
| 2021 | An exploratory analysis of regression methods for predicting faults in software systems
Santosh Singh Rathore |
Soft Comput. | 1 |
| 2021 | Generative Adversarial Networks-Based Imbalance Learning in Software Aging-Related Bug PredictionabstractSoftware aging refers to a problem of performance decay in the software systems, which are running for a long period. The primary cause of this phenomenon is the accumulation of run-time errors in the software, which are also known as aging-related bugs (ARBs). Many efforts have been reported earlier to predict the origin of ARBs in the software so that these bugs can be identified and fixed during testing. Imbalanced dataset, where the representation of ARBs patterns is very less as compared to the representation of the non-ARBs pattern significantly hinders the performance of the ARBs prediction models. Therefore, in this article, we present an oversampling approach, generative adversarial networks-based synthetic data generation-based ARBs prediction models. The approach uses generative adversarial networks to generate synthetic samples for the ARBs patterns in the given datasets implicitly and build the prediction models on the processed datasets. To validate the performance of the presented approach, we perform an experimental study for the seven ARBs datasets collected from the public repository and use various performance measures to evaluate the results. The experimental results showed that the presented approach led to the improved performance of prediction models for the ARBs prediction as compared to the other state-of-the-art models. Satyendra Singh Chouhan, Santosh Singh Rathore |
IEEE Trans. Reliab. | 2 |
| 2020 | Identifying Use Case Elements from Textual Specification: A Preliminary StudyabstractSoftware requirements are described in some form of natural language (NL) text so that stakeholders with limited experience can also comprehend them easily. However, the NL text written document is inherently ambiguous, and this makes it hard to examine requirements manually to find inconsistencies, duplicates, and/or missing requirements. Use Case Analysis is a graphical depiction used to explain the interaction between the user and the system for the given user's task. Additionally, it denotes the extension/dependency of one use case to another to understand the system flow. It is often used to identify, clarify, and categorize system requirements. However, generating use cases from a textual written description of requirements is an arduous task involving a significant manual work, which can be automated using data-driven techniques. In this poster paper, we present an initial approach for the automated identification of use case names and actor names from the textual requirements specification using machine learning techniques. Saurabh Tiwari 0001, Santosh Singh Rathore, Shreya Sagar, Yash Mirani |
RE | 2 |
| 2019 | Teaching Software Process Models to Software Engineering Students: An Exploratory StudyabstractA software process model (SPM) provides an abstract description of the order in which related activities of software development will be undertaken. Many process models available that can be adapted for software development. However, the selection of the best suitable process model with reference to the problem definition, constraints, and stakeholder requirements is a challenging task. Typically, in a Software Engineering (SE) course, students gain knowledge about SPM and realize their usage via classroom lectures and course projects. It is felt that if the basic knowledge imparted, through the fundamental SE course, is supplemented with some focused sessions about the SPM, then it will not only enable students to think in terms of the SPM but will also motivate them to harness the best practices of software development. This paper presents a preliminary study highlighting our experience on SPM-oriented teaching to impart the concept of requirement elicitation and process modeling, by performing a play (drama skit) annotating real-world scenarios. The feedbacks of students have been collected to evaluate whether this exercise helped them in understanding the processes they have to undergo during software development. Additionally, we have compared the student's feedback and performance in the project and reported the finding of the study. Saurabh Tiwari 0001, Santosh Singh Rathore |
APSEC | 2 |
| 2019 | An Approach for the Prediction of Number of Software Faults Based on the Dynamic Selection of Learning TechniquesabstractDetermining the most appropriate learning technique(s) is vital for the accurate and effective software fault prediction (SFP). Earlier techniques used for SFP have reported varying performance for different software projects and none of them has always reported the best performance across different projects. The problem of varying performance can be solved by using an approach, which partitions the fault dataset into different module subsets, trains learning techniques for each subset, and integrates the outcomes of all the learning techniques. This paper presents an approach that dynamically selects learning techniques to predict the number of software faults. For a given testing module, the presented approach first locates its neighbor module subset that contained modules similar to testing module using a distance function and then chooses the best learning technique in the region of that module subset to make the prediction for testing module. The learning technique is selected based on its past performance in the region of module subset. We have performed an evaluation of the proposed approach using fault datasets garnered from the PROMISE data repository and Eclipse bug data repository. Experimental results showed that the proposed approach led to an improved performance when predicting the number of faults in software systems. Santosh Singh Rathore, Sandeep Kumar 0004 |
IEEE Trans. Reliab. | 1 |
| 2017 | Towards an ensemble based system for predicting the number of software faults
Santosh Singh Rathore, Sandeep Kumar 0004 |
Expert Syst. Appl. | 1 |
| 2017 | Linear and non-linear heterogeneous ensemble methods to predict the number of faults in software systems
Santosh Singh Rathore, Sandeep Kumar 0004 |
Knowl. Based Syst. | 1 |
| 2017 | An empirical study of some software fault prediction techniques for the number of faults prediction
Santosh Singh Rathore, Sandeep Kumar 0004 |
Soft Comput. | 1 |
| 2012 | Validating the Effectiveness of Object-Oriented Metrics over Multiple Releases for Predicting Fault PronenessabstractIn this paper, we empirically investigate the re-lationship of existing class level object-oriented metrics with fault proneness over the multiple releases of the software. Here we first, evaluate each metric for their potential to predict faults independently by performing univariate logistic regression analysis. Next, we perform cross-correlation analysis between the significant metrics to find the subset of these metrics for an improved performance. The obtained metrics subset was then used to predict faults over the subsequent releases of the same project datasets. In this study, we used five publicly available project datasets over their multiple successive releases. Our results reported that the identified subset metrics demonstrated an improved fault prediction with higher accuracy and reduced misclassification errors. Santosh Singh Rathore, Atul Gupta |
APSEC | 1 |
| 2012 | An Approach to Generate Actor-Oriented Activity Charts from Use Case RequirementsabstractIn this paper, we propose an approach for transforming use case requirements into actor-oriented activity charts. In this approach, we first specify requirements using a relatively more formalized use case template. Next, we generate actor-oriented activity charts by identifying various interactions of an actor with the system and the events sequenced thereafter. The resulting activity charts can be useful in various software planning and development activities like release planning, test planning, execution and others. Saurabh Tiwari 0001, Santosh Singh Rathore, Abhijeet Singh, Atul Gupta |
APSEC | 2 |
| 2012 | Analysis of Use Case Requirements Using SFTA and SFMEA Techniques
Saurabh Tiwari 0001, Santosh Singh Rathore, Sudhanshu Gupta 0001, Gogate Vaibhav Vinayak, Atul Gupta |
ICECCS | 2 |