EDBT 2026 Demo / reviewers in the wild / expert
Anjana Gosain
dblp:72/2619
· DBLP profile ↗
19ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing medical data completeness using an iterative KNN based-Kernelized fuzzy c-means imputation methodabstractAccurate handling of missing values in medical datasets is essential to ensure reliable analysis, accurate diagnoses, and effective treatment planning. Incomplete medical data can significantly compromise the quality of clinical decision-making and hinder the development of intelligent healthcare systems. While numerous missing value imputation (MVI) methods have been proposed in the literature to address this challenge, most suffer from critical limitations such as difficulty in selecting optimal parameters and high sensitivity to noise and outliers. To address these limitations, this study introduces iKNN-KFCM, an iterative K-Nearest Neighbor-based Kernelized Fuzzy C-Means hybrid imputation method, specifically designed to enhance the quality and reliability of imputation in medical data. By synergistically integrating the local learning capability of K-Nearest Neighbors (KNN) with the non-linear clustering power of Kernelized Fuzzy C-Means (KFCM) within an iterative learning process, the proposed method refines missing value estimates by capturing the underlying structure and interrelationships within clinical datasets. Comprehensive experiments were conducted on five real-world medical datasets from the UCI repository with different levels of missing data to evaluate the effectiveness of our proposed iKNN-KFCM imputation method. The performance of this method was compared to eight other state-of-the-art imputation methods (mean, linear interpolation (LI), median, KNN imputation (KNNI), FCM imputation (FKMI), iterative fuzzy clustering (IFC), k-means imputation (KMI), and LI based FCM (LIFCM)) using several criteria such as root mean square error (RMSE), mean absolute error (MAE), mean imputation error (MIE), and mean square error (MSE). Additionally, Friedman’s statistical test was employed to validate the comparative results by ranking each imputation method based on its performance across different evaluation criteria. The experimental results indicated that the iKNN-KFCM method outperformed the existing methods across all evaluation criteria for these datasets, achieving superior performance in 94.29%, 88.57%, and 85.71% of instances for MIE, MAE, and RMSE, respectively, across all missing combinations. These results highlight the effectiveness and robustness of the iKNN-KFCM approach, reinforcing its suitability for improving the reliability of medical data analysis and supporting critical healthcare decision-making processes. Jyoti Singh, Jaspreeti Singh, Anjana Gosain |
Discov. Comput. | 3 |
| 2026 | Stack generalization-based hybrid ensemble classifier for imbalanced data
Suyash Kumar, Anjana Gosain |
Knowl. Inf. Syst. | 3 |
| 2025 | Kernelized Type 2 Intuitionistic Fuzzy C Means: a robust approach for brain MRI image segmentationabstractThis paper addresses the challenge of accurately segmenting complex regions in medical images, where traditional clustering methods often struggle due to noise sensitivity and unclear boundaries. Our objective is to develop a robust clustering approach for medical image segmentation. We present Kernelized Type 2 Intuitionistic Fuzzy C Means (KT2IFCM), which integrates a Radial Basis Function (RBF) kernel with Type 2 Intuitionistic Membership and hesitation degree. This method improves boundary definition, centroid placement, and handles non-linear structures. Results from tests on 14 datasets (4 simulated, 10 real brain MRI scans) show that KT2IFCM achieves superior noise resilience and segmentation accuracy compared to FCM, IFCM, KIFCM, and T2IFCM. The novelty lies in combining kernel mapping with intuitionistic fuzzy membership to deliver more reliable segmentation in medical imaging. Statistical analysis using the Friedman test further confirms KT2IFCM’s improved accuracy over competing methods on synthetic datasets. Kanika Bhalla, Sonika Dahiya, Anjana Gosain |
Discov. Comput. | 3 |
| 2025 | Hybrid resampling for enhanced multiclass classificationabstractMulticlass classification has become increasingly important in areas such as healthcare, fraud detection, agriculture, and network security. Yet, its effectiveness is often reduced when datasets are imbalanced, as conventional algorithms tend to favour majority classes while overlooking minority ones. Approaches like One-vs-One (OVO) and One-vs-All (OVA) attempt to manage this issue but often lead to information loss. Likewise, traditional resampling techniques either risk overfitting or remove useful data, limiting their impact in multiclass settings. The objective of this study is to identify a more effective strategy for handling multiclass imbalance by comparing a range of resampling methods and classifiers. We evaluated 14 well-known classifiers on 28 benchmark datasets from the UCI and KEEL repositories. Eight resampling methods, covering oversampling, undersampling, and hybrid techniques, were applied, with performance measured through cross-validation accuracy, accuracy, f1-score, and roc_auc. Our findings show that the hybrid resampling SmoteEnn combined with a Stacking Classifier consistently outperforms other methods, achieving up to 97.9% accuracy and f1-score, and 99.6% roc_auc. When benchmarked against the recently proposed Ensemble Partition Sampling (EPS), the proposed approach demonstrated clear superiority. These results confirm that hybrid resampling, particularly SmoteEnn, offers a robust and reliable solution for multiclass imbalanced learning. Suyash Kumar, Ritika Kumari, Anjana Gosain |
Discov. Comput. | 3 |
| 2025 | An effective imputation approach for handling missing data using intuitionistic fuzzy clustering algorithmsabstractIt is imperative to handle missing data attentively in the preprocessing stage as it may affects the integrity and quality of real-world datasets. However, existing soft clustering-based imputation neglect the underlying non-spherical separability of the data in feature space. This study proposes two robust missing data imputation (MDI) algorithms: Linear Interpolation-based Iterative Intuitionistic Fuzzy C-Means with Euclidean distance (LI-IIFCM) and its weighted variant LI-IIFCM-σ. LI-IIFCM and LI-IIFCM-σ uses linear interpolation for initial imputation followed by iterative IFCM and IFCM-σ, respectively. The approach leverages the soft Davies–Bouldin index to determine the optimal number of clusters and then iteratively refines imputations by minimizing average variation. Experimental analysis and statistical analysis (Friedman Test) on four UCI datasets, using two performance metrics, Mean Absolute Error (MAE) and Root Mean Square Error (RMSE), demonstrate that the proposed algorithms consistently outperform eight existing fuzzy clustering-based MDI algorithms. Kavita Sethia, Jaspreeti Singh, Anjana Gosain |
Discov. Comput. | 3 |
| 2023 | SmS: SMOTE-stacked hybrid model for diagnosis of polycystic ovary syndrome using feature selection method
Ritika Kumari, Jaspreeti Singh, Anjana Gosain |
Expert Syst. Appl. | 3 |
| 2023 | Prioritized dynamic cube selection in data warehouse
Heena Madaan, Anjana Gosain |
Multim. Tools Appl. | 2 |
| 2021 | Materialized view selection applying differential evolution algorithm combined with ensembled constraint handling techniques
Kavita Sachdeva, Anjana Gosain |
Multim. Tools Appl. | 2 |
| 2020 | WSEMQT: a novel approach for quality-based evaluation of web data sources for a data warehouseabstractThe incorporation of suitable external data from the World Wide Web offers an effective solution for enriching the data in the data warehouse (DW). However, the main challenge is the quality‐aware selection of web data sources to maintain the quality of the DW. In the previous works, the quality evaluation of web sources is through expert evaluation only, which makes it a very lengthy process. Also, since the quality model consists of mixed quality factors from diverse domains of Web, DW and underlying business, finding an expert possessing an expertise of all these domains is a huge bottleneck in the evaluation process. In order to overcome these existing issues, this study proposes a novel multi‐level approach web source evaluation with multi‐criteria decision‐making and web quality testing tools (WSEM QT ) and underlying quality model web quality model for evaluating web sources for the DW. The authors introduce automated web source quality evaluation in the first level of web source based evaluation and multiple dimensions of quality evaluation at the second level of expert‐based evaluation. At both the levels, multi‐criteria decision‐making methods are applied to the evaluation scores obtained to ascertain the ranked list of Web sources. The authors present a real‐world academic web data case study which shows that the proposed approach can be executed successfully for real‐world problems. Priyanka Bhutani, Anju Saha, Anjana Gosain |
IET Softw. | 3 |
| 2020 | Comprehensive complexity metric for data warehouse multidimensional model understandabilityabstractData warehouse quality can be determined during the initial phases of data warehouse development by quantifying the structural complexity of multidimensional models using metrics. The structural complexity of a multidimensional model is guided by its elements, types, and relationships among those elements. So far, most of the researchers have dealt with metrics based on various elements (facts, dimensions, dimensional hierarchies, and hierarchy levels) existing in these models. However, not much consideration is given to different types of dimensions based on hierarchy types and different relationships among those elements. Therefore, this work proposes a comprehensive complexity metric for measuring multidimensional model complexity by taking into account various elements, their types and the relationships among the elements at various levels of granularity in these models. The theoretical validation of the proposed metric using the property‐based framework given by Briand et al . characterises it as a complexity measure. Furthermore, the empirical study, employing statistical techniques (correlation and multinomial regression), on 26 multidimensional models and 20 subjects proved that the authors’ proposed metric is strongly correlated with multidimensional model understandability. Hence, this metric can be considered as a good predictor for data warehouse multidimensional model understandability. Anjana Gosain, Jaspreeti Singh |
IET Softw. | 1 |
| 2020 | A New Robust Fuzzy Clustering Approach: DBKIFCM
Anjana Gosain, Sonika Dahiya |
Neural Process. Lett. | 1 |
| 2020 | Robust hybrid data-level sampling approach to handle imbalanced data during classification
Anjana Gosain |
Soft Comput. | 2 |
| 2018 | Efficient approach for view materialisation in a data warehouse by prioritising data cubesabstractSelecting an appropriate set of views for materialisation is an important problem in a datawarehouse, and is referred to as the view selection problem. The existingstate‐of‐the‐art cost models select a set of views based on parameters, such asquery frequency, view size, view update frequency, and view update costs. Theexisting methods do not consider query priority as a parameter for selectingviews that can lead to shorter query processing times. Thus, in this paper,'priority’ is selected as a new selection parameter. Priority values areassigned to each query per user requirements, as well as using query type,user's level, and department preference in an organisation. As analyticalqueries require aggregated data cubes, priority values are assigned to each datacube based on priority value of the queries accessing them. Finally, a modifiedcost model is designed that integrates cube priority along with other selectionparameters. The authors’ proposed model uses the particle swarm optimisationalgorithm for selecting a set of prioritised cubes by minimising the total queryrunning cost under storage constraints. The experimental results shows that theproposed cost model leads to better cube selection, and consequently, shorterquery running times. Anjana Gosain, Heena Madaan |
IET Softw. | 1 |
| 2013 | Robust kernelized approach to clustering by incorporating new distance measure
A. K. Soni, Anjana Gosain |
Eng. Appl. Artif. Intell. | 3 |
| 2011 | A robust method for image segmentation of noisy digital imagesabstractA robust image segmentation algorithm called Extended Fuzzy C means (EFCM) is presented in this paper which preprocesses the image to reduce the noise effect and then apply FCM algorithm for image segmentation. Preprocessing of image is influenced by the direct eight neighborhood pixels of study pixel of an image under consideration. The advantages of the propose algorithm is: (1) Least execution time compared to other techniques. (2) It yields regions more homogeneous than those of other techniques. (3) It removes noisy spots and is less sensitive to noise. The propose technique is a powerful method for noisy image segmentation with least computation time and convergence rate compared to other image segmentation techniques. I. M. S. Lamba, Anjana Gosain |
FUZZ-IEEE | 3 |
| 2010 | Density-oriented approach to identify outliers and get noiseless clusters in Fuzzy C - MeansabstractIn an earlier work, we proposed Density Based Fuzzy C Means algorithm to identify noise and create clusters by changing Fuzzy C-Means (FCM) membership as well as objective functions. The constraint in changing membership in that algorithm produced a few unrealistic membership function values. In this paper, we propose Density Oriented Fuzzy C-Means (DOFCM) model that can detect efficient clusters in the presence of outliers and noise. DOFCM identifies outliers from a data-set before creating clusters and results into 'n+1' clusters, with 'n' good clusters and one invalid cluster containing noise and outliers. In this process, density approach has been used to identify outliers and modified FCM membership to create clusters. In DOFCM model, the location of the centroids is not affected by the presence of noise in the data-set. The results obtained through application of this model have been compared with various conventional and robust clustering techniques like FCM, PFCM, PCM, and NC, with the conclusion that the proposed technique gives better results. Anjana Gosain |
FUZZ-IEEE | 2 |
| 2009 | Improving the performance of fuzzy clustering algorithms through outlier identificationabstractMajor clustering algorithms consider all data objects as good objects while dividing data-set into clusters, except some, that consider noise/outliers to some extent. As a result those algorithms are not capable to produce efficient clusters as there is some effect of noise on location of cluster centroids. The task of outlier identification is to find small groups of data objects that are exceptional when compared with rest large amount of data. They are not required or acceptable while dividing a data-set into clusters, as clusters refer to the similar group of data and these outliers don't belong to any of the similar group. Yet they can be important in other applications. Through this paper we are trying to prove that efficient clusters can only be produced by identifying outliers and separating them from the data-set into one cluster before applying any clustering algorithm. In this paper a density based algorithm for outlier identification is proposed. Before applying any of the clustering algorithms; proposed algorithm is applied on the data-set to identify outliers and separate them from original data-set. Proposed algorithm is applied on fuzzy clustering algorithms (FCM, PCM and PFCM). Numerical examples and tests show that fuzzy algorithms after applying proposed algorithm gives better results when compared with the performance of fuzzy clustering algorithms without applying proposed technique. Anjana Gosain |
FUZZ-IEEE | 2 |
| 2008 | An approach to engineering the requirements of data warehouses
Naveen Prakash, Anjana Gosain |
Requir. Eng. | 2 |
| 2004 | Informational Scenarios for Data Warehouse Requirements Elicitation
Naveen Prakash, Yogesh Singh, Anjana Gosain |
ER | 3 |