VLDB 2026 Research / reviewers in the wild / expert
Ruchika Malhotra
dblp:09/4201
· DBLP profile ↗
35ranked-venue papers
30as first author
18since 2021 · last 2026
0000-0003-3872-6213ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 20 · 15 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 10 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DIR-SMOTE: a density-influence resampling framework for imbalanced code smell detection
Ruchika Malhotra, Bhawna Jain, Marouane Kessentini |
Autom. Softw. Eng. | 1 |
| 2026 | A Unified Deep Learning Based Feature Representation Approach for Effective Software Defect PredictionabstractABSTRACT Purpose Accurate software defect prediction (SDP) is critical to the success of any software project. Earlier studies have largely used static, semantic or structural features either in isolation or in pairs, offering a partial view of the source code. In reality, static features depict the statistical characteristics, semantic features depict the context and structural features depict data and control dependencies of the code. We propose a strong SDP model that integrates three types of features, achieving a holistic view of the source code, ultimately thereby enabling more robust and generalizable predictions. Methods First, the model extracts the static features from the open source PROMISE repository, semantic features from the Abstract Syntax Tree via CodeBERT followed by BiGRU and structural features from the Program Dependency Graph via Graph Convolutional Network. Second, feature alignment is performed for the fixed‐set representation of the three types of features using global attention pooling. Third, feature fusion is done for joint feature representation, followed by the application of additive attention to select the most suitable features. To handle the class imbalance scenario, cost‐sensitive gradient boosting is applied to penalize misclassifications more heavily. At last, the final feature set is fed to a classifier for defect prediction. Results Experiments conducted on eleven open source datasets reveal that the proposed unified feature representation approach achieves substantial performance improvements over the state‐of‐the‐art models. Moreover, the Wilcoxon signed‐rank test offers statistical validation for the relevance of these enhancements. Conclusion The integration of static, semantic and structural information results in a more holistic representation of source code, which substantially enhances defect prediction performance. The proposed approach addresses the partial view constraints of earlier approaches and offers strong potential for establishing a more reliable and robust SDP across various application domains. Ruchika Malhotra |
Softw. Pract. Exp. | 1 |
| 2025 | OpTunedSMOTE: A novel model for automated hyperparameter tuning of SMOTE in software defect predictionabstractSoftware Defect Prediction plays a crucial role in quality assurance by identifying potential defects early in the software development lifecycle. It is an essential aspect of modern software engineering that significantly contributes to improving software quality and reliability. It utilizes a variety of techniques, including machine learning algorithms like decision trees, support vector machines, and neural networks, to predict defects. A lot of research tried to improve the prediction accuracy but had problems with imbalanced data and hyperparameter tuning of the algorithms. To deal with this, we proposed a novel approach by tuning the hyperparameters of the Synthetic Minority Over-sampling Technique using the Tree-structured Parzen Estimator algorithm within the Optuna framework. Through an analysis of seventeen imbalanced datasets from a different public database, we compare our technique with existing SDP models using K-Nearest Neighbors, Multi-Layer Perceptron, Random Forest, Support Vector Machine, and Extreme Gradient Boosting classifiers. Our findings reveal that optimizing the Synthetic Minority Over-sampling Technique significantly improves the performance of SDP models, resulting in enhanced performance metrics. We have statistically validated our results using Friedman's test. Ruchika Malhotra, Kishwar Khan |
Intell. Data Anal. | 1 |
| 2025 | DHG-BiGRU: Dual-attention based hierarchical gated BiGRU for software defect prediction
Ruchika Malhotra |
Inf. Softw. Technol. | 1 |
| 2025 | Enhanced software change proneness prediction in android applications using balanced data techniques and advanced language modelsabstractSoftware change proneness prediction in Android applications is crucial in enhancing software maintenance and evolution processes. Inadequate consideration of data imbalance may compromise the reliability and robustness of change proneness prediction models. The study proposes a novel approach that leverages advanced techniques to address data imbalance while improving predictive accuracy. The methodology integrates state-of-the-art techniques to address key challenges. SMOTE (Synthetic Minority Over-sampling Technique) is employed for handling class imbalance, Self-Attention and RoBERTa embeddings are used to extract rich contextual information from text inputs, and Recursive Feature Elimination (RFE) is applied for feature selection to enhance model interpretability and performance. Moreover, we employ XLNet coupled with Butterfly optimization to reduce model weight while maintaining high predictive accuracy effectively. Our hybrid approach offers a robust framework for software change proneness prediction in Android applications by combining advanced language representation learning with optimization techniques. The efficiency of the presented technique is analyzed based on various metrics and research work compared with different state-of-the-art works. Ruchika Malhotra, Jyoti Patidar |
Discov. Comput. | 1 |
| 2025 | Software defect prediction based on multi-filter wrapper feature selection and deep neural network with attention mechanism
Ruchika Malhotra, Sonali Chawla |
Neural Comput. Appl. | 1 |
| 2025 | QNet: exploring deep learning for quantum code smell detection
Ruchika Malhotra, Bhawna Jain, Marouane Kessentini |
Softw. Qual. J. | 1 |
| 2025 | Improving software vulnerability severity prediction model performance with HDLN & FWFS: a two-stage feature selection approach
Ruchika Malhotra, Vidushi |
Softw. Qual. J. | 1 |
| 2024 | A systematic review of hyperparameter tuning techniques for software quality prediction modelsabstractBACKGROUND: Software quality prediction models play a crucial role in identifying vulnerable software components during early stages of development, and thereby optimizing the resource allocation and enhancing the overall software quality. While various classification algorithms have been employed for developing these prediction models, most studies have relied on default hyperparameter settings, leading to significant variability in model performance. Tuning the hyperparameters of classification algorithms can enhance the predictive capability of quality models by identifying optimal settings for improved accuracy and effectiveness. METHOD: This systematic review examines studies that have utilized hyperparameter tuning techniques to develop prediction models in software quality domain. The review focused on diverse areas such as defect prediction, maintenance estimation, change impact prediction, reliability prediction, and effort estimation, as these domains demonstrate the wide applicability of common learning algorithms. RESULTS: This review identified 31 primary studies on hyperparameter tuning for software quality prediction models. The results demonstrate that tuning the parameters of classification algorithms enhances the performance of prediction models. Additionally, the study found that certain classification algorithms exhibit high sensitivity to their parameter settings, achieving optimal performance when tuned appropriately. Conversely, certain classification algorithms exhibit low sensitivity to their parameter settings, making tuning unnecessary in such instances. CONCLUSION: Based on the findings of this review, the study conclude that the predictive capability of software quality prediction models can be significantly improved by tuning their hyperparameters. To facilitate effective hyperparameter tuning, we provide practical guidelines derived from the insights obtained through this study. Ruchika Malhotra, Madhukar Cherukuri |
Intell. Data Anal. | 1 |
| 2024 | A systematic review of transfer learning in software engineering
Ruchika Malhotra, Shweta Meena |
Multim. Tools Appl. | 1 |
| 2023 | Recent advances in deep learning models: a systematic literature review
Ruchika Malhotra |
Multim. Tools Appl. | 1 |
| 2023 | Software defect prediction using hybrid techniques: a systematic literature review
Ruchika Malhotra, Sonali Chawla |
Soft Comput. | 1 |
| 2022 | Handling class imbalance problem in software maintainability prediction: an empirical investigation
Ruchika Malhotra, Kusum Lata 0003 |
Frontiers Comput. Sci. | 1 |
| 2022 | Defect prediction model using transfer learning
Ruchika Malhotra, Shweta Meena |
Soft Comput. | 1 |
| 2021 | COVID'19 in India: Emotion of the Nation A novel Spatio-Temporal Unsupervised Sentiment AnalysisabstractThe coronavirus was originated in Wuhan City in China in 2019, and it led to something for which the world was not prepared. In the world where the extent of COVID’19 was ubiquitous, India was one of the countries that witnessed multiple phases of its spread. Moreover, since India has one of the largest populations, this made analyzing the sentiment of people during this time a task that held significance. COVID’19 brought a mix of emotions across the different periods in its first 18 months. During this time, social media was flooded with tweets and hashtags expressing both favorable and negative opinions about COVID’19, pandemic, lockdown, and vaccines. Ruchika Malhotra, Sarthak Aggarwal, Ridhima Bansal |
BIBM | 1 |
| 2021 | Comparative Study of Feature Reduction Techniques in Software Change PredictionabstractSoftware change prediction (SCP) is the process of identifying change-prone software classes using various structural and quality metrics by developing predictive techniques. The previous studies done in this field strongly confer the correlation between the quality of metrics and the performance of such SCP models. Past SCP studies have also applied different feature reduction (FR) techniques to address issues of high dimensionality, feature irrelevance, and feature repetition. Due to the vast variety of metric suites and FR techniques applied in SCP, there is a need to analyze and compare them. It will help in identifying the most crucial features and the most effective FR techniques. So, in this research, we conduct experiments to compare and contrast 60 Object-Oriented plus 26 Graph-based metrics and 11 state-of-the-art FR techniques previously employed for SCP over a range of 6 Java projects and 3 diverse classifiers. The AUC-ROC measures and statistical tests over experimental SCP models indicate that FR techniques are effective in SCP. Also, there exist significant differences in the performance of the different FR techniques. Furthermore, from this extensive experimentation, we were able to identify a set of the most effective FR techniques and the most crucial metrics which can be used to build effective SCP models. Ruchika Malhotra, Ritvik Kapoor, Deepti Aggarwal, Priya Garg |
MSR | 1 |
| 2021 | An empirical study to investigate the impact of data resampling techniques on the performance of class maintainability prediction models
Ruchika Malhotra, Kusum Lata 0003 |
Neurocomputing | 1 |
| 2021 | Predicting Software Defects for Object-Oriented Software Using Search-based TechniquesabstractDevelopment without any defect is unsubstantial. Timely detection of software defects favors the proper resource utilization saving time, effort and money. With the increasing size and complexity of software, demand for accurate and efficient prediction models is increasing. Recently, search-based techniques (SBTs) have fascinated many researchers for Software Defect Prediction (SDP). The goal of this study is to conduct an empirical evaluation to assess the applicability of SBTs for predicting software defects in object-oriented (OO) softwares. In this study, 16 SBTs are exploited to build defect prediction models for 13 OO software projects. Stable performance measures — GMean, Balance and Receiver Operating Characteristic-Area Under Curve (ROC-AUC) are employed to probe into the predictive capability of developed models, taking into consideration the imbalanced nature of software datasets. Proper measures are taken to handle the stochastic behavior of SBTs. The significance of results is statistically validated using the Friedman test complied with Wilcoxon post hoc analysis. The results confirm that software defects can be detected in the early phases of software development with help of SBTs. This paper identifies the effective subset of SBTs that will aid software practitioners to timely detect the probable software defects, therefore, saving resources and bringing up good quality softwares. Eight SBTs — sUpervised Classification System (UCS), Bioinformatics-oriented hierarchical evolutionary learning (BIOHEL), CHC, Genetic Algorithm-based Classifier System with Adaptive Discretization Intervals (GA_ADI), Genetic Algorithm-based Classifier System with Intervalar Rule (GA_INT), Memetic Pittsburgh Learning Classifier System (MPLCS), Population-Based Incremental Learning (PBIL) and Steady-State Genetic Algorithm for Instance Selection (SGA) are found to be statistically good defect predictors. Ruchika Malhotra, Juhi Jain |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2020 | Improving Software Maintainability Predictions using Data Oversampling and Hybridized TechniquesabstractSoftware systems being developed in today’s era are inherently large and complex. Maintaining these software is a big challenge before the software industry. Since software maintenance demands a high cost, this activity becomes even more challenging. In order to reduce the maintenance cost, it becomes crucial to know the maintainability of software modules/ classes in the initial stages of software development. Many efforts have been made to identify the maintainability of modules in the initial development stages in which prediction models have a lot of roles. Prediction models are trained from past historical data and should consist of an adequate number of instances of low maintainability and high maintainability class/modules. But this is not usually the case, and because of this, we are not able to train the prediction models properly. This situation shows data imbalance. In this direction, in this paper, we will handle the imbalanced data problem so that prediction models can be properly trained. We apply four oversampling techniques in this paper and train the prediction models for maintainability by hybridized techniques. The results of the paper advocate the effectiveness of examined oversampling techniques along with the hybridized classification techniques to develop competent maintainability prediction models. Ruchika Malhotra, Kusum Lata 0003 |
CEC | 1 |
| 2020 | A systematic literature review on empirical studies towards prediction of software maintainability
Ruchika Malhotra, Kusum Lata 0003 |
Soft Comput. | 1 |
| 2020 | An empirical study on predictability of software maintainability using imbalanced data
Ruchika Malhotra, Kusum Lata 0003 |
Softw. Qual. J. | 1 |
| 2019 | An empirical study to investigate oversampling methods for improving software defect prediction using imbalanced data
Ruchika Malhotra, Shine Kamal |
Neurocomputing | 1 |
| 2019 | Dynamic selection of fitness function for software change prediction using Particle Swarm Optimization
Ruchika Malhotra, Megha Khanna |
Inf. Softw. Technol. | 1 |
| 2018 | Threats to validity in search-based predictive modelling for software engineeringabstractA number of studies in the literature have developed effective models to address prediction tasks related to a software product such as estimating its development effort, or its change/defect proneness. These predictions are critical as they help in identifying weak areas of a software product and thus guide software project managers in effective allocation of project resources to these weak parts. Such practices assure good quality software products. Recently, the use of search‐based approaches (SBAs) for developing software prediction models (SPMs) has been successfully explored by a number of researchers. However, in order to develop effective and practical SPMs it is imperative to analyse various sources of threats. This study extensively reviews 93 primary studies, which use SBAs for developing SPMs of four commonly used software attributes (effort, defect‐proneness, maintainability and change‐proneness) in order to discuss and identify the various sources of threats while using these approaches for SPMs. The study also lists various actions that may be taken in order to minimise these threats. Furthermore, best practice examples in literature and the year‐wise trends of threats indicating the most common threats missed by researchers are provided to help academicians and practitioners in designing effective studies for developing SPMs using SBAs. Ruchika Malhotra, Megha Khanna |
IET Softw. | 1 |
| 2018 | Particle swarm optimization-based ensemble learning for software change prediction
Ruchika Malhotra, Megha Khanna |
Inf. Softw. Technol. | 1 |
| 2017 | An exploratory study for software change prediction in object-oriented systems using hybridized techniques
Ruchika Malhotra, Megha Khanna |
Autom. Softw. Eng. | 1 |
| 2017 | Special issue on search-based techniques and their hybridizations in software engineering
Ruchika Malhotra |
Comput. Lang. Syst. Struct. | 1 |
| 2017 | An empirical study for software change prediction using imbalanced data
Ruchika Malhotra, Megha Khanna |
Empir. Softw. Eng. | 1 |
| 2016 | Software Maintainability: Systematic Literature Review and Current TrendsabstractSoftware maintenance is an expensive activity that consumes a major portion of the cost of the total project. Various activities carried out during maintenance include the addition of new features, deletion of obsolete code, correction of errors, etc. Software maintainability means the ease with which these operations can be carried out. If the maintainability can be measured in early phases of the software development, it helps in better planning and optimum resource utilization. Measurement of design properties such as coupling, cohesion, etc. in early phases of development often leads us to derive the corresponding maintainability with the help of prediction models. In this paper, we performed a systematic review of the existing studies related to software maintainability from January 1991 to October 2015. In total, 96 primary studies were identified out of which 47 studies were from journals, 36 from conference proceedings and 13 from others. All studies were compiled in structured form and analyzed through numerous perspectives such as the use of design metrics, prediction model, tools, data sources, prediction accuracy, etc. According to the review results, we found that the use of machine learning algorithms in predicting maintainability has increased since 2005. The use of evolutionary algorithms has also begun in related sub-fields since 2010. We have observed that design metrics is still the most favored option to capture the characteristics of any given software before deploying it further in prediction model for determining the corresponding software maintainability. A significant increase in the use of public dataset for making the prediction models has also been observed and in this regard two public datasets User Interface Management System (UIMS) and Quality Evaluation System (QUES) proposed by Li and Henry is quite popular among researchers. Although machine learning algorithms are still the most popular methods, however, we suggest that researchers working on software maintainability area should experiment on the use of open source datasets with hybrid algorithms. In this regard, more empirical studies are also required to be conducted on a large number of datasets so that a generalized theory could be made. The current paper will be beneficial for practitioners, researchers and developers as they can use these models and metrics for creating benchmark and standards. Findings of this extensive review would also be useful for novices in the field of software maintainability as it not only provides explicit definitions, but also lays a foundation for further research by providing a quick link to all important studies in the said field. Finally, this study also compiles current trends, emerging sub-fields and identifies various opportunities of future research in the field of software maintainability. Ruchika Malhotra, Anuradha Chug |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2015 | Fault prediction considering threshold effects of object-oriented metricsabstractAbstract Software product quality can be enhanced significantly if we have a good knowledge and understanding of the potential faults therein. This paper describes a study to build predictive models to identify parts of the software that have high probability of occurrence of fault. We have considered the effect of thresholds of object‐oriented metrics on fault proneness and built predictive models based on the threshold values of the metrics used. Prediction of fault prone classes in earlier phases of software development life cycle will help software developers in allocating the resources efficiently. In this paper, we have used a statistical model derived from logistic regression to calculate the threshold values of object oriented, Chidamber and Kemerer metrics. Thresholds help developers to alarm the classes that fall outside a specified risk level. In this way, using the threshold values, we can divide the classes into two levels of risk – low risk and high risk. We have shown threshold effects at various risk levels and validated the use of these thresholds on a public domain, proprietary dataset, KC1 obtained from NASA and two open source, Promise datasets, IVY and JEdit using various machine learning methods and data mining classifiers. Interproject validation has also been carried out on three different open source datasets, Ant and Tomcat and Sakura. This will provide practitioners and researchers with well formed theories and generalised results. The results concluded that the proposed threshold methodology works well for the projects of similar nature or having similar characteristics. Ruchika Malhotra, Ankita Jain Bansal |
Expert Syst. J. Knowl. Eng. | 1 |
| 2010 | Empirical validation of object-oriented metrics for predicting fault proneness models
Yogesh Singh, Arvinder Kaur, Ruchika Malhotra |
Softw. Qual. J. | 3 |
| 2009 | Prediction of Software Quality Model Using Gene Expression Programming
Yogesh Singh, Arvinder Kaur, Ruchika Malhotra |
PROFES | 3 |
| 2008 | Predicting Software Fault Proneness Model Using Neural Network
Yogesh Singh, Arvinder Kaur, Ruchika Malhotra |
PROFES | 3 |
| 2008 | Predicting Software Fault Proneness Model Using Neural Network
Yogesh Singh, Arvinder Kaur, Ruchika Malhotra |
XP | 3 |
| 2005 | Software Reuse Metrics for Object-Oriented SystemsabstractThe importance of software measurement is increasing leading to development of new measurement techniques. Reusing existing software components is a key feature in increasing software productivity. It is one of the key elements in object-oriented programming, which reduces the cost and increases the quality of the software. An important feature of C++ called templates support generic programming, which allows the programmer to develop reusable software modules such as functions, classes, etc. The need for software reusability metrics is particularly acute for an organization in order to measure the degree of generic programming included in the form of templates in code. This research addresses this need and introduces a new set of metrics for object-oriented software. Two metrics are proposed for measuring amount of genericty included in the code and then analytically evaluated against Weyuker's set of nine axioms. This set of metrics is then applied to standard projects and accordingly ways in which project managers can use these metrics are suggested. K. K. Aggarwal 0001, Yogesh Singh, Arvinder Kaur, Ruchika Malhotra |
SERA | 4 |