Pradipta Maji

dblp:18/1028 · DBLP profile ↗
← Back
66ranked-venue papers
34as first author
13since 2021 · last 2024
0000-0002-8288-8917ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 8 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 12 · 8 first-author · 1 since 2021Theory of computation · 10 · 10 first-authorHuman-computer interaction and ubiquitous computing · 6 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection
Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Pradipta Maji, Hongmin Cai
ECCV (30)5
2024 Multi-Task Learning and Sparse Discriminant Canonical Correlation Analysis for Identification of Diagnosis-Specific Genotype-Phenotype Association
abstract
The primary objective of imaging genetics research is to investigate the complex genotype-phenotype association for the disease under study. For example, to understand the impact of genetic variations over the brain functions and structure, the genotypic data such as single nucleotide polymorphism (SNP) is integrated with the phenotypic data such as imaging quantitative traits. The sparse models, based on canonical correlation analysis (CCA), are popular in this area to find the complex bi-multivariate genotype-phenotype association, as the number of features in genotypic and/or phenotypic data is significantly higher as compared to the number of samples. However, the sparse CCA based methods are, in general, unsupervised in nature, and fail to identify the diagnose-specific features those play an important role for the diagnosis and prognosis of the disease under study. In this regard, a new supervised model is proposed to study the complex genotype-phenotype association, by judiciously integrating the merits of CCA, linear discriminant analysis (LDA) and multi-task learning. The proposed model can identify the diagnose-specific as well as the diagnose-consistent features with significantly lower computational complexity. The performance of the proposed method, along with a comparison with the state-of-the-art methods, is evaluated on several synthetic data sets and one real imaging genetics data collected from Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort. In the current study, the SNP as genetic data and resting state functional MRI ( fMRI) as imaging data are integrated to find the complex genotype-phenotype association. An important finding is that the proposed method has better correlation value, improved noise resistance and stability, and also has better feature selection ability. All the results illustrate the power and capability of the proposed method to find the diagnostic group-specific imaging genetic association, which may help to understand the neurodegenerative disorder in a more comprehensive way.
Sankar Mondal, Pradipta Maji
IEEE ACM Trans. Comput. Biol. Bioinform.2
2024 Discriminative Deep Canonical Correlation Analysis for Multi-View Data
abstract
Over the past few years, multimodal data analysis has emerged as an inevitable method for identifying sample categories. In the multi-view data classification problem, it is expected that the joint representation should include the supervised information of sample categories so that the similarity in the latent space implies the similarity in the corresponding concepts. Since each view has different statistical properties, the joint representation should be able to encapsulate the underlying nonlinear data distribution of the given observations. Another important aspect is the coherent knowledge of the multiple views. It is required that the learning objective of the multi-view model efficiently captures the nonlinear correlated structures across different modalities. In this context, this article introduces a novel architecture, termed discriminative deep canonical correlation analysis (D2CCA), for classifying given observations into multiple categories. The learning objective of the proposed architecture includes the merits of generative models to identify the underlying probability distribution of the given observations. In order to improve the discriminative ability of the proposed architecture, the supervised information is incorporated into the learning objective of the proposed model. It also enables the architecture to serve as both a feature extractor as well as a classifier. The theory of CCA is integrated with the objective function so that the joint representation of the multi-view data is learned from maximally correlated subspaces. The proposed framework is consolidated with corresponding convergence analysis. The efficacy of the proposed architecture is studied on different domains of applications, namely, object recognition, document classification, multilingual categorization, face recognition, and cancer subtype identification with reference to several state-of-the-art methods.
Debamita Kumar, Pradipta Maji
IEEE Trans. Neural Networks Learn. Syst.2
2023 Multi-View Kernel Learning for Identification of Disease Genes
abstract
Gene expression data sets and protein-protein interaction (PPI) networks are two heterogeneous data sources that have been extensively studied, due to their ability to capture the co-expression patterns among genes and their topological connections. Although they depict different traits of the data, both of them tend to group co-functional genes together. This phenomenon agrees with the basic assumption of multi-view kernel learning, according to which different views of the data contain a similar inherent cluster structure. Based on this inference, a new multi-view kernel learning based disease gene identification algorithm, termed as DiGId, is put forward. A novel multi-view kernel learning approach is proposed that aims to learn a consensus kernel, which efficiently captures the heterogeneous information of individual views as well as depicts the underlying inherent cluster structure. Some low-rank constraints are imposed on the learned multi-view kernel, so that it can effectively be partitioned into k or fewer clusters. The learned joint cluster structure is used to curate a set of potential disease genes. Moreover, a novel approach is put forward to quantify the importance of each view. In order to demonstrate the effectiveness of the proposed approach in capturing the relevant information depicted by individual views, an extensive analysis is performed on four different cancer-related gene expression data sets and PPI network, considering different similarity measures.
Ekta Shah, Pradipta Maji
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Multiview Regularized Discriminant Canonical Correlation Analysis: Sequential Extraction of Relevant Features From Multiblock Data
abstract
One of the important issues associated with real-life high-dimensional data analysis is how to extract significant and relevant features from multiview data. The multiset canonical correlation analysis (MCCA) is a well-known statistical method for multiview data integration. It finds a linear subspace that maximizes the correlations among different views. However, the existing methods to find the multiset canonical variables are computationally very expensive, which restricts the application of the MCCA in real-life big data analysis. The covariance matrix of each high-dimensional view may also suffer from the singularity problem due to the limited number of samples. Moreover, the MCCA-based existing feature extraction algorithms are, in general, unsupervised in nature. In this regard, a new supervised feature extraction algorithm is proposed, which integrates multimodal multidimensional data sets by solving maximal correlation problem of the MCCA. A new block matrix representation is introduced to reduce the computational complexity for computing the canonical variables of the MCCA. The analytical formulation enables efficient computation of the multiset canonical variables under supervised ridge regression optimization technique. It deals with the "curse of dimensionality" problem associated with high-dimensional data and facilitates the sequential generation of relevant features with significantly lower computational cost. The effectiveness of the proposed multiblock data integration algorithm, along with a comparison with other existing methods, is demonstrated on several benchmark and real-life cancer data.
Ankita Mandal, Pradipta Maji
IEEE Trans. Cybern.2
2023 Adaptive Generalized Multi-View Canonical Correlation Analysis for Incrementally Update Multiblock Data
abstract
One of the major problems in real-life multiblock dynamic data analysis is that all the available modalities may not be relevant. Some of them may provide noisy or even inconsistent information with respect to other modalities. So, it is necessary to evaluate the quality of a new modality before considering it for feature extraction. In this regard, the paper introduces a new multiset canonical correlation analysis (MCCA), termed as incremental MCCA (IMCCA). When a new modality is available for the analysis, the IMCCA generates the new canonical variables from that of the earlier modalities, without repeating the same procedure with the original data augmented by the new modality. The proposed IMCCA deals with the “curse of dimensionality” problem associated with multidimensional data sets, by using the ridge regression optimization technique. Using the proposed IMCCA model, a new feature extraction algorithm is introduced, which considers a new modality for the analysis if it has relevant and significant information with respect to existing modalities. The proposed algorithm starts with the two most relevant modalities, and the remaining modalities are added sequentially according to their relevance. The optimum regularization parameters for the proposed algorithm are estimated based on the supervised information of sample categories. The effectiveness of the proposed algorithm, along with a comparison with state-of-the-art multimodal data integration methods, is established on several real-life multiblock data sets.
Ankita Mandal, Pradipta Maji
IEEE Trans. Knowl. Data Eng.2
2023 Truncated Normal Mixture Prior Based Deep Latent Model for Color Normalization of Histology Images
abstract
The variation in color appearance among the Hematoxylin and Eosin (H&E) stained histological images is one of the major problems, as the color disagreement may affect the computer aided diagnosis of histology slides. In this regard, the paper introduces a new deep generative model to reduce the color variation present among the histological images. The proposed model assumes that the latent color appearance information, extracted through a color appearance encoder, and stain bound information, extracted via stain density encoder, are independent of each other. In order to capture the disentangled color appearance and stain bound information, a generative module as well as a reconstructive module are considered in the proposed model to formulate the corresponding objective functions. The discriminator is modeled to discriminate between not only the image samples, but also the joint distributions corresponding to image samples, color appearance information and stain bound information, which are sampled individually from different source distributions. To deal with the overlapping nature of histochemical reagents, the proposed model assumes that the latent color appearance code is sampled from a mixture model. As the outer tails of a mixture model do not contribute adequately in handling overlapping information, rather are prone to outliers, a mixture of truncated normal distributions is used to deal with the overlapping nature of histochemical stains. The performance of the proposed model, along with a comparison with state-of-the-art approaches, is demonstrated on several publicly available data sets containing H&E stained histological images. An important finding is that the proposed model outperforms state-of-the-art methods in 91.67% and 69.05% cases, with respect to stain separation and color normalization, respectively.
Suman Mahapatra, Pradipta Maji
IEEE Trans. Medical Imaging2
2022 Scalable Non-Linear Graph Fusion for Prioritizing Cancer-Causing Genes
abstract
In the past few decades, both gene expression data and protein-protein interaction (PPI)networks have been extensively studied, due to their ability to depict important characteristics of disease-associated genes. In this regard, the paper presents a new gene prioritization algorithm to identify and prioritize cancer-causing genes, integrating judiciously the complementary information obtained from two data sources. The proposed algorithm selects disease-causing genes by maximizing the importance of selected genes and functional similarity among them. A new quantitative index is introduced to evaluate the importance of a gene. It considers whether a gene exhibits a differential expression pattern across sick and healthy individuals, and has a strong connectivity in the PPI network, which are the important characteristics of a potential biomarker. As disease-associated genes are expected to have similar expression profiles and topological structures, a scalable non-linear graph fusion technique, termed as ScaNGraF, is proposed to learn a disease-dependent functional similarity network from the co-expression and common neighbor based similarity networks. The proposed ScaNGraF, which is based on message passing algorithm, efficiently combines the shared and complementary information provided by different data sources with significantly lower computational cost. A new measure, termed as DiCoIN, is introduced to evaluate the quality of a learned affinity network. The performance of the proposed graph fusion technique and gene selection algorithm is extensively compared with that of some existing methods, using several cancer data sets.
Ekta Shah, Pradipta Maji
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 Selective Update of Relevant Eigenspaces for Integrative Clustering of Multimodal Data
abstract
One of the major problems in cancer subtype discovery from multimodal omic data is that all the available modalities may not encode relevant and homogeneous information about the subtypes. Moreover, the high-dimensional nature of the modalities makes sample clustering computationally expensive. In this regard, a novel algorithm is proposed to extract a low-rank joint subspace of the integrated data matrix. The proposed algorithm first evaluates the quality of subtype information provided by each of the modalities, and then judiciously selects only relevant ones to construct the joint subspace. The problem of incrementally updating the singular value decomposition of a data matrix is formulated for the multimodal data framework. The analytical formulation enables efficient construction of the joint subspace of integrated data from low-rank subspaces of the individual modalities. The construction of joint subspace by the proposed method is shown to be computationally more efficient compared to performing the principal component analysis (PCA) on the integrated data matrix. Some new quantitative indices are introduced to measure theoretically the accuracy of subspace construction by the proposed approach with respect to the principal subspace extracted by the PCA. The efficacy of clustering on the joint subspace constructed by the proposed algorithm is established over existing integrative clustering approaches on several real-life multimodal cancer data sets.
Aparajita Khan, Pradipta Maji
IEEE Trans. Cybern.2
2022 Multi-Manifold Optimization for Multi-View Subspace Clustering
abstract
The meaningful patterns embedded in high-dimensional multi-view data sets typically tend to have a much more compact representation that often lies close to a low-dimensional manifold. Identification of hidden structures in such data mainly depends on the proper modeling of the geometry of low-dimensional manifolds. In this regard, this article presents a manifold optimization-based integrative clustering algorithm for multi-view data. To identify consensus clusters, the algorithm constructs a joint graph Laplacian that contains denoised cluster information of the individual views. It optimizes a joint clustering objective while reducing the disagreement between the cluster structures conveyed by the joint and individual views. The optimization is performed alternatively over k -means and Stiefel manifolds. The Stiefel manifold helps to model the nonlinearities and differential clusters within the individual views, whereas k -means manifold tries to elucidate the best-fit joint cluster structure of the data. A gradient-based movement is performed separately on the manifold of each view so that individual nonlinearity is preserved while looking for shared cluster information. The convergence of the proposed algorithm is established over the manifold and asymptotic convergence bound is obtained to quantify theoretically how fast the sequence of iterates generated by the algorithm converges to an optimal solution. The integrative clustering on benchmark and multi-omics cancer data sets demonstrates that the proposed algorithm outperforms state-of-the-art multi-view clustering approaches.
Aparajita Khan, Pradipta Maji
IEEE Trans. Neural Networks Learn. Syst.2
2021 Approximate Graph Laplacians for Multimodal Data Clustering
abstract
One of the important approaches of handling data heterogeneity in multimodal data clustering is modeling each modality using a separate similarity graph. Information from the multiple graphs is integrated by combining them into a unified graph. A major challenge here is how to preserve cluster information while removing noise from individual graphs. In this regard, a novel algorithm, termed as CoALa, is proposed that integrates noise-free approximations of multiple similarity graphs. The proposed method first approximates a graph using the most informative eigenpairs of its Laplacian which contain cluster information. The approximate Laplacians are then integrated for the construction of a low-rank subspace that best preserves overall cluster information of multiple graphs. However, this approximate subspace differs from the full-rank subspace which integrates information from all the eigenpairs of each Laplacian. Matrix perturbation theory is used to theoretically evaluate how far approximate subspace deviates from the full-rank one for a given value of approximation rank. Finally, spectral clustering is performed on the approximate subspace to identify the clusters. Experimental results on several real-life cancer and benchmark data sets demonstrate that the proposed algorithm significantly and consistently outperforms state-of-the-art integrative clustering approaches.
Aparajita Khan, Pradipta Maji
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Rough-Bayesian approach to select class-pair specific descriptors for HEp-2 cell staining pattern recognition
Debamita Kumar, Pradipta Maji
Pattern Recognit.2
2021 Rough Hypercuboid Based Generalized and Robust IT2 Fuzzy C-Means Algorithm
abstract
One of the important issues in pattern recognition and machine learning is how to find natural groups present in a dataset. In this regard, this paper presents a novel clustering algorithm, called rough hypercuboid-based interval type-2 fuzzy c -means (RIT2FCM). It judiciously integrates the merits of the rough hypercuboid approach, c -means algorithm, and interval type-2 fuzzy set, to address the uncertainty associated with real-life datasets. Using the concept of hypercuboid equivalence partition matrix (HEM) of rough hypercuboid approach, the lower approximation and boundary region of each cluster are implicitly defined, without using any prespecified threshold parameter. The interval-valued fuzzifier is applied to address the uncertainty coupled with different parameters of rough-fuzzy clustering algorithms, where the determination of the appropriate value of fuzzifier is a difficult task. An analytical formulation on the convergence analysis of the proposed RIT2FCM algorithm, along with a theoretical bound of its fuzzifier, is also introduced. The efficacy of the proposed RIT2FCM method is extensively compared with that of several existing clustering algorithms, using some cluster validity and classification rate indices on various real-life datasets. The proposed algorithm performs better than the state-of-the-art c -means algorithms in 92.59% cases, with respect to different cluster validity indices, in lesser computation time.
Pradipta Maji, Partha Garai
IEEE Trans. Cybern.1
2020 CanSuR: a robust method for staining pattern recognition of HEp-2 cell IIF images
Ankita Mandal, Pradipta Maji
Neural Comput. Appl.2
2020 Rough segmentation of coherent local intensity for bias induced 3-D MR brain images
Shaswati Roy, Pradipta Maji
Pattern Recognit.2
2020 Low-Rank Joint Subspace Construction for Cancer Subtype Discovery
abstract
Multimodal data integration is an important framework for cancer subtype discovery as it can blend the inherent properties of individual modalities with their cross-platform correlations to infer clinically relevant subtypes. The main problem here is the appropriate selection of relevant and complementary modalities. Another problem is the 'high dimension-low sample size' nature of each modality. The current research work proposes a novel algorithm to construct a low-rank joint subspace from the low-rank subspaces of individual high-dimensional modalities. Statistical hypothesis testing is introduced to effectively estimate the rank of each modality by separating the signal component from its noise counterpart. Two quantitative indices are proposed to evaluate the quality of different modalities, the first one assesses the degree of relevance of the cluster structure embedded within each modality, while the second measure evaluates the amount of cluster information shared between two modalities. To construct the joint subspace, the algorithm selects the most relevant modalities with maximum shared information. During data integration, the intersection between two subspaces is also considered to select cluster information and filter out the noise from different subspaces. The efficacy of clustering on the joint subspace, extracted by the proposed algorithm, is compared with that of several existing integrative clustering approaches on real-life multimodal data sets. Experimental results show that the identified subtypes have closer resemblance with the clinically established subtypes as compared to the subtypes identified by the existing approaches. Survival analysis has revealed the significant differences between survival profiles of the identified subtypes, while robustness analysis shows that the identified subtypes are not sensitive towards perturbation of the data sets.
Aparajita Khan, Pradipta Maji
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 Medical Image Segmentation by Partitioning Spatially Constrained Fuzzy Approximation Spaces
abstract
Image segmentation is an important prerequisite step for any automatic clinical analysis technique. It assists in visualization of human tissues, as accurate delineation of medical images requires involvement of expert practitioners, which is also time consuming. In this background, the rough-fuzzy clustering algorithm provides an effective approach for image segmentation. It handles uncertainties arising due to overlapping classes and incompleteness in class definition by partitioning the fuzzy approximation spaces. However, the existing rough-fuzzy clustering algorithms do not consider the spatial distribution of the image. They depend only on the distribution of pixels to determine their class labels. In this regard, this article introduces a new algorithm, termed as spatially constrained rough-fuzzy c-means (sRFCM) for medical image segmentation. The proposed sRFCM algorithm combines wisely the merits of rough-fuzzy clustering and local neighborhood information. In the proposed algorithm, the labels of local neighbors influence in the determination of the label of center pixel. The effect of local neighbors acts as a regularizer. Moreover, the proposed sRFCM algorithm partitions each cluster in possibilistic lower approximation or core region and probabilistic boundary region. The cluster centroid depends on the core and boundary regions, weight parameter, and neighborhood regularizer. A novel segmentation validity index, termed as neighborhood Silhouette, is proposed to find out the optimum values of regularizer and weight parameter, controlling the performance of the sRFCM. The efficacy of the proposed sRFCM algorithm, as well as several existing segmentation algorithms, is demonstrated on four brain MR volume databases and one HEp-2 cell image data.
Shaswati Roy, Pradipta Maji
IEEE Trans. Fuzzy Syst.2
2020 A Spatially Constrained Probabilistic Model for Robust Image Segmentation
abstract
In general, the hidden Markov random field (HMRF) represents the class label distribution of an image in probabilistic model based segmentation. The class label distributions provided by existing HMRF models consider either the number of neighboring pixels with similar class labels or the spatial distance of neighboring pixels with dissimilar class labels. Also, this spatial information is only considered for estimation of class labels of the image pixels, while its contribution in parameter estimation is completely ignored. This, in turn, deteriorates the parameter estimation, resulting in sub-optimal segmentation performance. Moreover, the existing models assign equal weightage to the spatial information for class label estimation of all pixels throughout the image, which, create significant misclassification for the pixels in boundary region of image classes. In this regard, the paper develops a new clique potential function and a new class label distribution, incorporating the information of image class parameters. Unlike existing HMRF model based segmentation techniques, the proposed framework introduces a new scaling parameter that adaptively measures the contribution of spatial information for class label estimation of image pixels. The importance of the proposed framework is depicted by modifying the HMRF based segmentation methods. The advantage of proposed class label distribution is also demonstrated irrespective of the underlying intensity distributions. The comparative performance of the proposed and existing class label distributions in HMRF model is demonstrated both qualitatively and quantitatively for brain MR image segmentation, HEp-2 cell delineation, natural image and object segmentation.
Abhirup Banerjee, Pradipta Maji
IEEE Trans. Image Process.2
2020 Circular Clustering in Fuzzy Approximation Spaces for Color Normalization of Histological Images
abstract
One of the foremost and challenging tasks in hematoxylin and eosin stained histological image analysis is to reduce color variation present among images, which may significantly affect the performance of computer-aided histological image analysis. In this regard, the paper introduces a new rough-fuzzy circular clustering algorithm for stain color normalization. It judiciously integrates the merits of both fuzzy and rough sets. While the theory of rough sets deals with uncertainty, vagueness, and incompleteness in stain class definition, fuzzy set handles the overlapping nature of histochemical stains. The proposed circular clustering algorithm works on a weighted hue histogram, which considers both saturation and local neighborhood information of the given image. A new dissimilarity measure is introduced to deal with the circular nature of hue values. Some new quantitative measures are also proposed to evaluate the color constancy after normalization. The performance of the proposed method, along with a comparison with other state-of-the-art methods, is demonstrated on several publicly available standard data sets consisting of hematoxylin and eosin stained histological images.
Pradipta Maji, Suman Mahapatra
IEEE Trans. Medical Imaging1
2019 Rough-Fuzzy Circular Clustering for Color Normalization of Histological Images
abstract
Color disagreement among histological images may affect the performance of computer-aided histological image analysis. So, one of the most important and challenging tasks in histological image analysis is to diminish the color variation among the images, maintaining the histological information contained in them. In this regard, the paper proposes a new circular clustering algorithm, termed as rough-fuzzy circular clustering. It integrates judiciously the merits of rough-fuzzy clustering and cosine distance. The rough-fuzzy circular clustering addresses the uncertainty due to vagueness and incompleteness in stain class definition, as well as overlapping nature of multiple contrasting histochemical stains. The proposed circular clustering algorithm incorporates saturation-weighted hue histogram, which considers both saturation and hue information of the given histological image. The efficacy of the proposed method, along with a comparison with other state-of-the-art methods, is demonstrated on publicly available hematoxylin and eosin stained fifty-eight benchmark histological images.
Pradipta Maji, Suman Mahapatra
Fundam. Informaticae1
2019 Segmentation of bias field induced brain MR images using rough sets and stomped-t distribution
Abhirup Banerjee, Pradipta Maji
Inf. Sci.2
2018 FaRoC: Fast and Robust Supervised Canonical Correlation Analysis for Multimodal Omics Data
abstract
One of the main problems associated with high dimensional multimodal real life data sets is how to extract relevant and significant features. In this regard, a fast and robust feature extraction algorithm, termed as FaRoC, is proposed, integrating judiciously the merits of canonical correlation analysis (CCA) and rough sets. The proposed method extracts new features sequentially from two multidimensional data sets by maximizing their relevance with respect to class label and significance with respect to already-extracted features. To generate canonical variables sequentially, an analytical formulation is introduced to establish the relation between regularization parameters and CCA. The formulation enables the proposed method to extract required number of correlated features sequentially with lesser computational cost as compared to existing methods. To compute both significance and relevance measures of a feature, the concept of hypercuboid equivalence partition matrix of rough hypercuboid approach is used. It also provides an efficient way to find optimum regularization parameters employed in CCA. The efficacy of the proposed FaRoC algorithm, along with a comparison with other existing methods, is extensively established on several real life data sets.
Ankita Mandal, Pradipta Maji
IEEE Trans. Cybern.2
2017 Rough Hypercuboid and Modified Kulczynski Coefficient for Disease Gene Identification
Ekta Shah, Pradipta Maji
ACIIDS (2)2
2017 Stomped-t: A novel probability distribution for rough-probabilistic clustering
Abhirup Banerjee, Pradipta Maji
Inf. Sci.2
2017 RelSim: An integrated method to identify disease genes using gene expression profiles and PPIN based similarity measure
Pradipta Maji, Ekta Shah, Sushmita Paul
Inf. Sci.1
2017 Significance and Functional Similarity for Identification of Disease Genes
abstract
One of the most significant research issues in functional genomics is insilico identification of disease related genes. In this regard, the paper presents a new gene selection algorithm, termed as SiFS, for identification of disease genes. It integrates the information obtained from interaction network of proteins and gene expression profiles. The proposed SiFS algorithm culls out a subset of genes from microarray data as disease genes by maximizing both significance and functional similarity of the selected gene subset. Based on the gene expression profiles, the significance of a gene with respect to another gene is computed using mutual information. On the other hand, a new measure of similarity is introduced to compute the functional similarity between two genes. Information derived from the protein-protein interaction network forms the basis of the proposed SiFS algorithm. The performance of the proposed gene selection algorithm and new similarity measure, is compared with that of other related methods and similarity measures, using several cancer microarray data sets.
Pradipta Maji, Ekta Shah
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 A modified rough-fuzzy clustering algorithm with spatial information for HEp-2 cell image segmentation
abstract
Indirect immunofluorescence (IIF) analysis is the most effective test for antinuclear autoantibodies (ANAs) analysis, in order to reveal the occurrence of some autoimmune diseases, such as connective tissue disorders. In the tests of antinuclear antibodies, the human epithelial type 2 (HEp-2) cells is mostly used as substrate. However, the recognition of the staining pattern of ANAs in the IIF image requires proper detection of the region of interest. In this regard, automatic segmentation of IIF images is an essential prerequisite as manual segmentation is labor intensive, time consuming, and subjective. Recently, rough-fuzzy clustering has been shown to provide significant results for image segmentation by handling different uncertainties present in the images. But, the existing robust rough-fuzzy clustering algorithm does not consider spatial distribution of the image. This is useful when the image is distorted by noise and other artifacts. In this regard, the paper proposes a segmentation algorithm by incorporating the spatial constraint with the advantages of robust rough-fuzzy clustering. In the current study, class label of a pixel is influenced by its neighboring pixels depending on their spatial distance. In this way, more number of neighboring pixels can be incorporated into the calculation of a pixel feature. The performance of the proposed method is evaluated on several HEp-2 cell images and compared with the existing algorithms by presenting both qualitative and quantitative results.
Shaswati Roy, Pradipta Maji
BIBM2
2016 Rough Hypercuboid Based Supervised Regularized Canonical Correlation for Multimodal Data Analysis
abstract
One of the main problems in real life omics data analysis is how to extract relevant and non-redundant features from high dimensional multimodal data sets. In general, supervised regularized canonical correlation analysis (SRCCA) plays an important role in extracting new features from multimodal om ics data sets. However, the existing SRCCA optimizes regularization parameters based on the quality of first pair of canonical variables only using standard feature evaluation indices. In this regard, this paper introduces a new SRCCA algorithm, integrating judiciously the merits of SRCCA and rough hypercuboid approach, to extract relevant and non-redundant features in approximation spaces from multimodal omics data sets. The proposed method optimizes regularization parameters of the SRCCA based on the quality of a set of pairs of canonical variables using rough hypercuboid approach. While the rough hypercuboid approach provides an efficient way to calculate the degree of dependency of class labels on feature set in approximation spaces, the merit of SRCCA helps in extracting non-redundant features from multimodal data sets. The effectiveness of the proposed approach, along with a comparison with related existing approaches, is demonstrated on several real life data sets.
Pradipta Maji, Ankita Mandal
Fundam. Informaticae1
2016 Preface: pattern recognition and mining
Pradipta Maji, Sankar K. Pal, Andrzej Skowron
Nat. Comput.1
2016 Gene expression and protein-protein interaction data for identification of colon cancer related genes using f-information measures
Sushmita Paul, Pradipta Maji
Nat. Comput.2
2015 Simultaneous Feature Selection and Extraction Using Feature Significance
abstract
Dimensionality reduction of a data set by selecting or extracting relevant and nonredundant features is an essential preprocessing step used for pattern recognition, data mining, machine learning, and multimedia indexing. Among the large amount of features present in real life data sets, only a small fraction of them is effective to represent the data set accurately. Prior to analysis of the data set, preprocessing the data to obtain a smaller set of representative features and retaining the optimal salient characteristics of the data not only decrease the processing time but also lead to more compactness of the models learned and better generalization. In this regard, a novel dimensionality reduction method is presented here that simultaneously selects and extracts features using the concept of feature significance. The method is based on maximizing both relevance and significance of the reduced feature set, whereby redundancy therein is removed. The method is generic in nature in the sense that both supervised and unsupervised feature evaluation indices can be used for simultaneously feature selection and extraction. The effectiveness of the proposed method, along with a comparison with existing feature selection and extraction methods, is demonstrated on a set of real life data sets.
Pradipta Maji, Partha Garai
Fundam. Informaticae1
2015 SoBT-RFW: Rough-Fuzzy Computing and Wavelet Analysis Based Automatic Brain Tumor Detection Method from MR Images
abstract
One of the important problems in medical diagnosis is the segmentation and detection of brain tumor in MR images. The accurate estimation of brain tumor size is important for treatment planning and therapy evaluation. In this regard, this paper prese
Pradipta Maji, Shaswati Roy
Fundam. Informaticae1
2015 IT2 Fuzzy-Rough Sets and Max Relevance-Max Significance Criterion for Attribute Selection
abstract
One of the important problems in pattern recognition, machine learning, and data mining is the dimensionality reduction by attribute or feature selection. In this regard, this paper presents a feature selection method, based on interval type-2 (IT2) fuzzy-rough sets, where the features are selected by maximizing both relevance and significance of the features. By introducing the concept of lower and upper fuzzy equivalence partition matrices, the lower and upper relevance and significance of the features are defined for IT2 fuzzy approximation spaces. Different feature evaluation criteria such as dependency, relevance, and significance are presented for attribute selection task using IT2 fuzzy-rough sets. The performance of IT2 fuzzy-rough sets is compared with that of some existing feature evaluation indices including classical rough sets, neighborhood rough sets, and type-1 fuzzy-rough sets. The effectiveness of the proposed IT2 fuzzy-rough set-based attribute selection method, along with a comparison with existing feature selection and extraction methods, is demonstrated on several real-life data.
Pradipta Maji, Partha Garai
IEEE Trans. Cybern.1
2015 Rough Sets and Stomped Normal Distribution for Simultaneous Segmentation and Bias Field Correction in Brain MR Images
abstract
The segmentation of brain MR images into different tissue classes is an important task for automatic image analysis technique, particularly due to the presence of intensity inhomogeneity artifact in MR images. In this regard, this paper presents a novel approach for simultaneous segmentation and bias field correction in brain MR images. It integrates judiciously the concept of rough sets and the merit of a novel probability distribution, called stomped normal (SN) distribution. The intensity distribution of a tissue class is represented by SN distribution, where each tissue class consists of a crisp lower approximation and a probabilistic boundary region. The intensity distribution of brain MR image is modeled as a mixture of finite number of SN distributions and one uniform distribution. The proposed method incorporates both the expectation-maximization and hidden Markov random field frameworks to provide an accurate and robust segmentation. The performance of the proposed approach, along with a comparison with related methods, is demonstrated on a set of synthetic and real brain MR images for different bias fields and noise levels.
Abhirup Banerjee, Pradipta Maji
IEEE Trans. Image Process.2
2014 A Rough Hypercuboid Approach for Feature Selection in Approximation Spaces
abstract
The selection of relevant and significant features is an important problem particularly for data sets with large number of features. In this regard, a new feature selection algorithm is presented based on a rough hypercuboid approach. It selects a set of features from a data set by maximizing the relevance, dependency, and significance of the selected features. By introducing the concept of the hypercuboid equivalence partition matrix, a novel representation of degree of dependency of sample categories on features is proposed to measure the relevance, dependency, and significance of features in approximation spaces. The equivalence partition matrix also offers an efficient way to calculate many more quantitative measures to describe the inexactness of approximate classification. Several quantitative indices are introduced based on the rough hypercuboid approach for evaluating the performance of the proposed method. The superiority of the proposed method over other feature selection methods, in terms of computational complexity and classification accuracy, is established extensively on various real-life data sets of different sizes and dimensions.
Pradipta Maji
IEEE Trans. Knowl. Data Eng.1
2013 Contraharmonic Mean Based Bias Field Correction in MR Images
Abhirup Banerjee, Pradipta Maji
CAIP (1)2
2013 µHEM for identification of differentially expressed miRNAs using hypercuboid equivalence partition matrix
abstract
BACKGROUND: The miRNAs, a class of short approximately 22-nucleotide non-coding RNAs, often act post-transcriptionally to inhibit mRNA expression. In effect, they control gene expression by targeting mRNA. They also help in carrying out normal functioning of a cell as they play an important role in various cellular processes. However, dysregulation of miRNAs is found to be a major cause of a disease. It has been demonstrated that miRNA expression is altered in many human cancers, suggesting that they may play an important role as disease biomarkers. Multiple reports have also noted the utility of miRNAs for the diagnosis of cancer. Among the large number of miRNAs present in a microarray data, a modest number might be sufficient to classify human cancers. Hence, the identification of differentially expressed miRNAs is an important problem particularly for the data sets with large number of miRNAs and small number of samples. RESULTS: In this regard, a new miRNA selection algorithm, called μHEM, is presented based on rough hypercuboid approach. It selects a set of miRNAs from a microarray data by maximizing both relevance and significance of the selected miRNAs. The degree of dependency of sample categories on miRNAs is defined, based on the concept of hypercuboid equivalence partition matrix, to measure both relevance and significance of miRNAs. The effectiveness of the new approach is demonstrated on six publicly available miRNA expression data sets using support vector machine. The.632+ bootstrap error estimate is used to minimize the variability and biasedness of the derived results. CONCLUSIONS: An important finding is that the μHEM algorithm achieves lowest B.632+ error rate of support vector machine with a reduced set of differentially expressed miRNAs on four expression data sets compare to some existing machine learning and statistical methods, while for other two data sets, the error rate of the μHEM algorithm is comparable with the existing techniques. The results on several microarray data sets demonstrate that the proposed method can bring a remarkable improvement on miRNA selection problem. The method is a potentially useful tool for exploration of miRNA expression data and identification of differentially expressed miRNAs worth further investigation.
Sushmita Paul, Pradipta Maji
BMC Bioinform.2
2013 Robust Rough-Fuzzy C-Means Algorithm: Design and Applications in Coding and Non-coding RNA Expression Data Clustering
abstract
Cluster analysis is a technique that divides a given data set into a set of clusters in such a way that two objects from the same cluster are as similar as possible and the objects from different clusters are as dissimilar as possible. In this backgr
Pradipta Maji, Sushmita Paul
Fundam. Informaticae1
2013 Rough-Fuzzy Clustering for Grouping Functionally Similar Genes from Microarray Data
abstract
Gene expression data clustering is one of the important tasks of functional genomics as it provides a powerful tool for studying functional relationships of genes in a biological process. Identifying coexpressed groups of genes represents the basic challenge in gene clustering problem. In this regard, a gene clustering algorithm, termed as robust rough-fuzzy c-means, is proposed judiciously integrating the merits of rough sets and fuzzy sets. While the concept of lower and upper approximations of rough sets deals with uncertainty, vagueness, and incompleteness in cluster definition, the integration of probabilistic and possibilistic memberships of fuzzy sets enables efficient handling of overlapping partitions in noisy environment. The concept of possibilistic lower bound and probabilistic boundary of a cluster, introduced in robust rough-fuzzy c-means, enables efficient selection of gene clusters. An efficient method is proposed to select initial prototypes of different gene clusters, which enables the proposed c-means algorithm to converge to an optimum or near optimum solutions and helps to discover coexpressed gene clusters. The effectiveness of the algorithm, along with a comparison with other algorithms, is demonstrated both qualitatively and quantitatively on 14 yeast microarray data sets.
Pradipta Maji, Sushmita Paul
IEEE ACM Trans. Comput. Biol. Bioinform.1
2013 Fuzzy-Rough Simultaneous Attribute Selection and Feature Extraction Algorithm
abstract
Among the huge number of attributes or features present in real-life data sets, only a small fraction of them are effective to represent the data set accurately. Prior to analysis of the data set, selecting or extracting relevant and significant features is an important preprocessing step used for pattern recognition, data mining, and machine learning. In this regard, a novel dimensionality reduction method, based on fuzzy-rough sets, that simultaneously selects attributes and extracts features using the concept of feature significance is presented. The method is based on maximizing both the relevance and significance of the reduced feature set, whereby redundancy therein is removed. This paper also presents classical and neighborhood rough sets for computing the relevance and significance of the feature set and compares their performances with that of fuzzy-rough sets based on the predictive accuracy of nearest neighbor rule, support vector machine, and decision tree. An important finding is that the proposed dimensionality reduction method based on fuzzy-rough sets is shown to be more effective for generating a relevant and significant feature subset. The effectiveness of the proposed fuzzy-rough-set-based dimensionality reduction method, along with a comparison with existing attribute selection and feature extraction methods, is demonstrated on real-life data sets.
Pradipta Maji, Partha Garai
IEEE Trans. Cybern.1
2013 Rough Sets for Bias Field Correction in MR Images Using Contraharmonic Mean and Quantitative Index
abstract
One of the challenging tasks for magnetic resonance (MR) image analysis is to remove the intensity inhomogeneity artifact present in MR images, which often degrades the performance of an automatic image analysis technique. In this regard, the paper presents a novel approach for bias field correction in MR images. It judiciously integrates the merits of rough sets and contraharmonic mean. While the contraharmonic mean is used in low-pass averaging filter to estimate the bias field in multiplicative model, the concept of lower approximation and boundary region of rough sets deals with vagueness and incompleteness in filter structure definition. A theoretical analysis is presented to justify the use of both rough sets and contraharmonic mean for bias field estimation. The integration enables the algorithm to estimate optimum or near optimum bias field. Some new quantitative indexes are introduced to measure intensity inhomogeneity artifact present in a MR image. The performance of the proposed approach, along with a comparison with other approaches, is demonstrated on both simulated and real MR images for different bias fields and noise levels.
Abhirup Banerjee, Pradipta Maji
IEEE Trans. Medical Imaging2
2012 Robust RFCM algorithm for identification of co-expressed miRNAs
abstract
MicroRNAs (miRNAs) are short, endogenous RNAs having ability to regulate gene expression at the post-transcriptional level. Various studies have revealed that miRNAs tend to cluster on chromosomes. Members of a cluster that are at close proximity on chromosome are highly likely to be processed as cotranscribed units. Therefore, a large proportion of miRNAs are co-expressed. Expression profiling of miRNAs generates a huge volume of data. Complicated networks of miRNA-mRNA interaction create a big challenge for scientists to decipher this huge expression data. In order to extract meaningful information from expression data, this paper presents the application of robust rough-fuzzy c-means (rRFCM) algorithm to discover co-expressed miRNA clusters. The rRFCM algorithm comprises a judicious integration of rough sets, fuzzy sets, and c-means algorithm. The effectiveness of the rRFCM algorithm and different initialization methods, along with a comparison with other related methods, is demonstrated on three miRNA microarray expression data sets using Silhouette index, Davies-Bouldin index, Dunn index, β index, and gene ontology based analysis.
Sushmita Paul, Pradipta Maji
BIBM2
2012 Fuzzy-Rough MRMS Method for Relevant and Significant Attribute Selection
Pradipta Maji, Partha Garai
IPMU (1)1
2012 Mutual Information-Based Supervised Attribute Clustering for Microarray Sample Classification
abstract
Microarray technology is one of the important biotechnological means that allows to record the expression levels of thousands of genes simultaneously within a number of different samples. An important application of microarray gene expression data in functional genomics is to classify samples according to their gene expression profiles. Among the large amount of genes presented in gene expression data, only a small fraction of them is effective for performing a certain diagnostic test. Hence, one of the major tasks with the gene expression data is to find groups of coregulated genes whose collective expression is strongly associated with the sample categories or response variables. In this regard, a new supervised attribute clustering algorithm is proposed to find such groups of genes. It directly incorporates the information of sample categories into the attribute clustering process. A new quantitative measure, based on mutual information, is introduced that incorporates the information of sample categories to measure the similarity between attributes. The proposed supervised attribute clustering algorithm is based on measuring the similarity between attributes using the new quantitative measure, whereby redundancy among the attributes is removed. The clusters are then refined incrementally based on sample categories. The performance of the proposed algorithm is compared with that of existing supervised and unsupervised gene clustering and gene selection algorithms based on the class separability index and the predictive accuracy of naive bayes classifier, K-nearest neighbor rule, and support vector machine on three cancer and two arthritis microarray data sets. The biological significance of the generated clusters is interpreted using the gene ontology. An important finding is that the proposed supervised attribute clustering algorithm is shown to be effective for identifying biologically significant gene clusters with excellent predictive capability.
Pradipta Maji
IEEE Trans. Knowl. Data Eng.1
2011 Microarray Time-Series Data Clustering Using Rough-Fuzzy C-Means Algorithm
abstract
Clustering is one of the important analysis in functional genomics that discovers groups of co-expressed genes from microarray data. In this paper, the application of rough-fuzzy c-means (RFCM) algorithm is presented to discover co-expressed gene clusters. One of the major issues of the RFCM based microarray data clustering is how to select initial prototypes of different clusters. To overcome this limitation, a method is proposed to select initial cluster centers. It enables the RFCM algorithm to converge to an optimum or near optimum solutions and helps to discover co-expressed gene clusters. A method is also introduced based on Dunn's cluster validity index to identify optimum values of different parameters of the initialization method and the RFCM algorithm. The effectiveness of the RFCM algorithm, along with a comparison with other related methods, is demonstrated on five yeast gene expression time-series data sets using Silhouette index, Davies-Bouldin index, and gene ontology based analysis.
Pradipta Maji, Sushmita Paul
BIBM1
2011 Rough set based maximum relevance-maximum significance criterion and Gene selection from microarray data
Pradipta Maji, Sushmita Paul
Int. J. Approx. Reason.1
2011 Fuzzy-Rough Supervised Attribute Clustering Algorithm and Classification of Microarray Data
abstract
One of the major tasks with gene expression data is to find groups of coregulated genes whose collective expression is strongly associated with sample categories. In this regard, a new clustering algorithm, termed as fuzzy-rough supervised attribute clustering (FRSAC), is proposed to find such groups of genes. The proposed algorithm is based on the theory of fuzzy-rough sets, which directly incorporates the information of sample categories into the gene clustering process. A new quantitative measure is introduced based on fuzzy-rough sets that incorporates the information of sample categories to measure the similarity among genes. The proposed algorithm is based on measuring the similarity between genes using the new quantitative measure, whereby redundancy among the genes is removed. The clusters are refined incrementally based on sample categories. The effectiveness of the proposed FRSAC algorithm, along with a comparison with existing supervised and unsupervised gene selection and clustering algorithms, is demonstrated on six cancer and two arthritis data sets based on the class separability index and predictive accuracy of the naive Bayes' classifier, the K-nearest neighbor rule, and the support vector machine.
Pradipta Maji
IEEE Trans. Syst. Man Cybern. Part B1
2010 Feature Selection Using f-Information Measures in Fuzzy Approximation Spaces
abstract
The selection of nonredundant and relevant features of real-valued data sets is a highly challenging problem. A novel feature selection method is presented here based on fuzzy-rough sets by maximizing the relevance and minimizing the redundancy of the selected features. By introducing the fuzzy equivalence partition matrix, a novel representation of Shannon's entropy for fuzzy approximation spaces is proposed to measure the relevance and redundancy of features suitable for real-valued data sets. The fuzzy equivalence partition matrix also offers an efficient way to calculate many more information measures, termed as f-information measures. Several f-information measures are shown to be effective for selecting nonredundant and relevant features of real-valued data sets. This paper compares the performance of different f-information measures for feature selection in fuzzy approximation spaces. Some quantitative indexes are introduced based on fuzzy-rough sets for evaluating the performance of proposed method. The effectiveness of the proposed method, along with a comparison with other methods, is demonstrated on a set of real-life data sets.
Pradipta Maji, Sankar K. Pal
IEEE Trans. Knowl. Data Eng.1
2010 Fuzzy-Rough Sets for Information Measures and Selection of Relevant Genes From Microarray Data
abstract
Several information measures such as entropy, mutual information, and f-information have been shown to be successful for selecting a set of relevant and nonredundant genes from a high-dimensional microarray data set. However, for continuous gene expression values, it is very difficult to find the true density functions and to perform the integrations required to compute different information measures. In this regard, the concept of the fuzzy equivalence partition matrix is presented to approximate the true marginal and joint distributions of continuous gene expression values. The fuzzy equivalence partition matrix is based on the theory of fuzzy-rough sets, where each row of the matrix represents a fuzzy equivalence partition that can automatically be derived from the given expression values. The performance of the proposed approach is compared with that of existing approaches using the class separability index and the predictive accuracy of the support vector machine. An important finding, however, is that the proposed approach is shown to be effective for selecting relevant and nonredundant continuous-valued genes from microarray data.
Pradipta Maji, Sankar K. Pal
IEEE Trans. Syst. Man Cybern. Part B1
2010 Rough Sets for Selection of Molecular Descriptors to Predict Biological Activity of Molecules
abstract
Quantitative structure activity relationship (QSAR) is one of the important disciplines of computer-aided drug design that deals with the predictive modeling of properties of a molecule. In general, each QSAR dataset is small in size with large number of features or descriptors. Among the large amount of descriptors presented in the QSAR dataset, only a small fraction of them is effective for performing the predictive modeling task. In this paper, a new feature selection algorithm is presented, based on rough set theory, to select a set of effective molecular descriptors from a given QSAR dataset. The proposed algorithm selects the set of molecular descriptors by maximizing both relevance and significance of the descriptors. An important finding is that the proposed feature selection algorithm is shown to be effective in selecting relevant and significant molecular descriptors from the QSAR dataset for predictive modeling. The performance of the proposed algorithm is studied using R2statistic of support vector regression method. The effectiveness of the proposed algorithm, along with a comparison with existing algorithms, is demonstrated on three QSAR datasets.
Pradipta Maji, Sushmita Paul
IEEE Trans. Syst. Man Cybern. Part C1
2009 Content-based image retrieval using visually significant point features
Minakshi Banerjee, Malay Kumar Kundu, Pradipta Maji
Fuzzy Sets Syst.3
2008 On Characterization of Attractor Basins of Fuzzy Multiple Attractor Cellular Automata
Pradipta Maji
Fundam. Informaticae1
2008 Second Order Fuzzy Measure and Weighted Co-Occurrence Matrix for Segmentation of Brain MR Images
Pradipta Maji, Malay Kumar Kundu, Bhabatosh Chanda
Fundam. Informaticae1
2008 Efficient design of neural network tree using a new splitting criterion
Pradipta Maji
Neurocomputing1
2008 Non-uniform cellular automata based associative memory: Evolutionary design and basins of attraction
Pradipta Maji, Parimal Pal Chaudhuri
Inf. Sci.1
2007 RBFFCA: A Hybrid Pattern Classifier Using Radial Basis Function and Fuzzy Cellular Automata
Pradipta Maji, Parimal Pal Chaudhuri
Fundam. Informaticae1
2007 RFCM: A Hybrid Clustering Algorithm Using Rough and Fuzzy Sets
Pradipta Maji, Sankar K. Pal
Fundam. Informaticae1
2007 Rough-Fuzzy C-Medoids Algorithm and Selection of Bio-Basis for Amino Acid Sequence Analysis
abstract
In most pattern recognition algorithms, amino acids cannot be used directly as inputs since they are nonnumerical variables. They, therefore, need encoding prior to input. In this regard, bio-basis function maps a nonnumerical sequence space to a numerical feature space. It is designed using an amino acid mutation matrix. One of the important issues for the bio-basis function is how to select the minimum set of bio-bases with maximum information. In this paper, we describe an algorithm, termed as rough-fuzzy c{\hbox{-}}{\rm{medoids}} (RFCMdd) algorithm, to select the most informative bio-bases. It is comprised of a judicious integration of the principles of rough sets, fuzzy sets, the c{\hbox{-}}{\rm{medoids}} algorithm, and the amino acid mutation matrix. While the membership function of fuzzy sets enables efficient handling of overlapping partitions, the concept of lower and upper bounds of rough sets deals with uncertainty, vagueness, and incompleteness in class definition. The concept of crisp lower bound and fuzzy boundary of a class, introduced in RFCMdd, enables efficient selection of the minimum set of the most informative bio-bases. Some new indices are introduced for evaluating quantitatively the quality of selected bio-bases. The effectiveness of the proposed algorithm, along with a comparison with other algorithms, has been demonstrated on different types of protein data sets.
Pradipta Maji, Sankar K. Pal
IEEE Trans. Knowl. Data Eng.1
2007 Rough Set Based Generalized Fuzzy C-Means Algorithm and Quantitative Indices
abstract
A generalized hybrid unsupervised learning algorithm, which is termed as rough-fuzzy possibilistic c-means (RFPCM), is proposed in this paper. It comprises a judicious integration of the principles of rough and fuzzy sets. While the concept of lower and upper approximations of rough sets deals with uncertainty, vagueness, and incompleteness in class definition, the membership function of fuzzy sets enables efficient handling of overlapping partitions. It incorporates both probabilistic and possibilistic memberships simultaneously to avoid the problems of noise sensitivity of fuzzy c-means and the coincident clusters of PCM. The concept of crisp lower bound and fuzzy boundary of a class, which is introduced in the RFPCM, enables efficient selection of cluster prototypes. The algorithm is generalized in the sense that all existing variants of c-means algorithms can be derived from the proposed algorithm as a special case. Several quantitative indices are introduced based on rough sets for the evaluation of performance of the proposed c-means algorithm. The effectiveness of the algorithm, along with a comparison with other algorithms, has been demonstrated both qualitatively and quantitatively on a set of real-life data sets.
Pradipta Maji, Sankar K. Pal
IEEE Trans. Syst. Man Cybern. Part B1
2004 FMACA: A Fuzzy Cellular Automata Based Pattern Classifier
Pradipta Maji, Parimal Pal Chaudhuri
DASFAA1
2004 Cellular Automata Based Pattern Classifying Machine for Distributed Data Mining
Pradipta Maji, Parimal Pal Chaudhuri
ICONIP1
2004 Design and characterization of cellular automata based associative memory for pattern recognition
abstract
This paper reports a cellular automata (CA) based model of associative memory. The model has been evolved around a special class of CA referred to as generalized multiple attractor cellular automata (GMACA). The GMACA based associative memory is designed to address the problem of pattern recognition. Its storage capacity is found to be better than that of Hopfield network. The GMACA are configured with nonlinear CA rules that are evolved through genetic algorithm (GA). Successive generations of GA select the rules at the edge of chaos. The study confirms the potential of GMACA to perform complex computations like pattern recognition at the edge of chaos.
Niloy Ganguly, Pradipta Maji, Biplab K. Sikdar, Parimal Pal Chaudhuri
IEEE Trans. Syst. Man Cybern. Part B2
2003 Theory and Application of Cellular Automata For Pattern Classification
Pradipta Maji, Chandrama Shaw, Niloy Ganguly, Biplab K. Sikdar, Parimal Pal Chaudhuri
Fundam. Informaticae1
2003 Error correcting capability of cellular automata based associative memory
abstract
This paper reports the error correcting capability of an associative memory model built around the sparse network of cellular automata (CA). Analytical formulation supported by experimental results has demonstrated the capability of CA based sparse network to memorize unbiased patterns while accommodating noise. The desired CA are evolved with an efficient formulation of simulated annealing (SA) program. The simple, regular, modular, and cascadable structure of CA based associative memory suits ideally for design of low cost high speed online pattern recognizing machine with the currently available VLSI technology.
Pradipta Maji, Niloy Ganguly, Parimal Pal Chaudhuri
IEEE Trans. Syst. Man Cybern. Part A1
2002 Generalized Multiple Attractor Cellular Automata (GMACA) Model for Associative Memory
abstract
This paper reports an efficient technique of evolving Cellular Automata (CA) as an associative memory model. The evolved CA termed as GMACA (Generalized Multiple Attractor Cellular Automata), acts as a powerful pattern recognizer. Detailed analysis of GMACA rules establishes the fact that the rule subspace of the pattern recognizing CA lies at the edge of chaos — believed to be capable of executing complex computation.
Niloy Ganguly, Pradipta Maji, Biplab K. Sikdar, Parimal Pal Chaudhuri
Int. J. Pattern Recognit. Artif. Intell.2
2001 Evolving Cellular Automata Based Associative Memory for Pattern Recognition
Niloy Ganguly, Arijit Das, Pradipta Maji, Biplab K. Sikdar, Parimal Pal Chaudhuri
HiPC3