VLDB 2026 Research / reviewers in the wild / expert
Sanghamitra Bandyopadhyay
dblp:b/SanghamitraBandyopadhyay
· DBLP profile ↗
149ranked-venue papers
46as first author
25since 2021 · last 2026
0000-0001-6370-2083ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 19 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 48 · 12 first-author · 11 since 2021Databases, data management, data science and information retrieval · 18 · 8 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 14 · 5 first-author · 1 since 2021Theory of computation · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Attention-Enhanced Architecture with Multi-objective Hyperparameter Optimization for Efficient Lung Segmentation
Sudipta Kumar Biswas, Mrittika Chakraborty, Rangan Das, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
ACIIDS (2) | 5 |
| 2026 | BKDRP: a biological knowledge-driven approach for drug response prediction using multi-omics data in cancer cell linesabstractBACKGROUND: Cancer heterogeneity results in patients with the same diagnosis responding differently to drugs, making treatments extremely challenging. Advances in computational power enable personalized treatments that suppress tumors and extend patient survival. Therefore, accurate prediction of cancer cell response to a particular medication is of utmost importance. Current deep learning-based models have achieved impressive accuracy, but they often function as a "black box" and cannot explain the reason for the prediction. To address this limitation, we develop a deep learning-based model, BKDRP, which incorporates prior biological information into the architecture, along with molecular fingerprints of drugs, while embedding biological priors into its architecture. Specifically, it incorporates the fact that genes encode proteins that combine to form protein complexes, which in turn regulate biological pathways, ultimately targeted by drugs. RESULTS: We evaluate BKDRP on the GDSC (Genomics of Drug Sensitivity and Cancer) cell line dataset using multi-omics gene expression, protein expression, mutation, and copy number variation. Four rigorous experiments have been conducted to test the model's generalizability: prediction of unknown drug-cell line responses, responses to unseen drugs (LODO: Leave-One-Drug-Out), responses to unseen cell lines (LOCLO: Leave-On-Cell-Line-Out), and responses across unseen cancer types (LOCO: Leave-One-Cancer-Out). The performance of the proposed method and baseline algorithms is assessed using two metrics: Area Under the ROC Curve (AUC) and Area Under the Precision-Recall Curve (AUPR). The experimental results demonstrate that BKDRP performs well in different evaluation techniques. Notably, BKDRP has achieved an AUC of 0.8845, surpassing traditional machine learning and deep learning approaches and demonstrating robustness in handling biological variability across cancer types. A case study of lung adenocarcinoma (LUAD) highlights known biomarkers (KRAS, EGFR, STK11), key proteins (SOCS1, HSPA8, SMC3), and drugs (Erlotinib, Palbociclib) that are consistent with the literature. CONCLUSIONS: In conclusion, BKDRP presents a novel biological knowledge-driven deep neural network model for cancer drug response prediction that shows strong predictive accuracy and interpretability. By integrating multi-omics data and incorporating domain knowledge, BKDRP has the strong potential for applications in biomarker discovery and the advancement of personalized oncology. Koyel Mandal, Sanghamitra Bandyopadhyay |
BMC Bioinform. | 2 |
| 2026 | Drug Effect Classification Using Frequency-Based Graph Traversal ApproachabstractClassifying drugs into symptomatic (SYM) and disease-modifying (DM) categories is essential for understanding their therapeutic effect and plays a key role in drug repurposing. While many computational approaches focus on drug-target prediction, they often ignore the nature of the drug's action on disease progression. This study proposes a graph-based strategy to classify drugs as SYM or DM based on their effect on disease treatment. We construct a heterogeneous network comprising genes, diseases, and drugs, and apply a guided shortest path traversal framework for drug effect classification. During this traversal, certain genes appear frequently in the shortest metapaths linking diseases and drugs. These recurrent genes are identified based on their frequency of occurrence in known drug-disease paths. For a new drug-disease pair, if the traversal path contains recurrent genes marked for a specific treatment type, we classify the drug accordingly. Over and above classifying the drugs, the proposed method incorporates the metapath-based framework to improve interpretability. Experimental results show that our model achieves significantly better classification accuracy compared to advanced machine learning and deep learning methods. A case study on multiple sclerosis further supports the biological relevance of our approach. Aishik Chanda, Ashmita Dey, Mrittika Chakraborty, Utsav Bandyopadhyay Maulik, Sanghamitra Bandyopadhyay |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | GroupEx: Toward Group-Level Explanations of Graph Neural NetworksabstractThe increasing adoption of graph neural networks (GNNs) in critical real-world applications brings with it theneed for interpretability and explainability of these models. The existing paradigms of GNN explainability typically operate at the extremes. Instance-level methods offer a microscopic view focused on individual predictions, whereas model-level approaches provide a macroscopic summary of the classifier behavior on a target class. We introduce the group-level paradigm, which views explanations at the right resolution, grounded in the intuition that groups of instances embedded close together in the latent space share a common rationale for belonging to the same target class. We show that the instance-level and model-level paradigms are the special cases of this broader framework. To realize this paradigm, we propose GroupEx, a novel method for interpreting GNNs at the group level. GroupEx has the capability to identify the subgroups of graphs within a target class using a group identity (GI) score. It employs a group extractor that can extract a subgroup of graphs that share a common explanation for their class identity by clustering graph embeddings in the latent space of the classifier. A specially devised explanation generator can then discover a group-level explanation by optimizing a novel isomorphism objective, which ensures that the discovered explanation commonly explains the class identity of all the instances in the group. Empirical results on real-world and synthetic datasets demonstrate that GroupEx provides deeper insights into GNN decision-making than the existing state-of-the-art explainability methods, enabling a more structured and interpretable understanding of model predictions. Sayan Saha 0001, Sanghamitra Bandyopadhyay |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Enhancing Single-Cell RNA-Seq Data Completeness With a Graph Learning FrameworkabstractSingle cell RNA sequencing (scRNA-seq) is a powerful tool to capture gene expression snapshots in individual cells. However, a low amount of RNA in the individual cells results in dropout events, which introduce huge zero counts in the single cell expression matrix. We have developed VAImpute, a variational graph autoencoder based imputation technique that learns the inherent distribution of a large network/graph constructed from the scRNA-seq data leveraging copula correlation ($Ccor$) among cells/genes. The trained model is utilized to predict the dropouts events by computing the probability of all non-edges (cell-gene) in the network. We devise an algorithm to impute the missing expression values of the detected dropouts. The performance of the proposed model is assessed on both simulated and real scRNA-seq datasets, comparing it to established single-cell imputation methods. VAImpute yields significant improvements to detect dropouts, thereby achieving superior performance in cell clustering, detecting rare cells, and differential expression. Snehalika Lall, Sumanta Ray, Sanghamitra Bandyopadhyay |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Gen-GraphEx: Generative In-Distribution Graph Explanations for Time-Efficient Model-Level Interpretability of GNNsabstractGraph neural networks (GNNs) have become the prevailing methodology for addressing graph data-related tasks, permeating critical domains like recommendation systems and drug development. The necessity for trustworthiness and interpretability of GNNs has risen to the forefront, especially given their direct impact on end users' lives. To address this need, we present Gen-GraphEx, a model-agnostic, model-level explanation method that prioritizes user centricness by eliminating the need for having access to the hidden layers of the GNN model it seeks to explain. Given a particular class label, Gen-GraphEx learns a graph generative model (GGM) that produces explanation graphs that not only contain discriminative patterns that the GNN has learned for that class but also lie in distribution with real graphs that belong to that class according to the GNN. Unlike existing state-of-the-art models, Gen-GraphEx also has the unique ability to interpolate the GGMs of two target classes to generate instances that lie near the decision boundary of the two classes giving a deeper insight into the model's decision-making. Its advantages over existing methods in the literature also include nonreliance on another subsequent deep learning module for explanation generation, ability to generate graphs with various node and edge features, and being more computationally efficient. Extensive validation and thorough comparative analysis of the proposed approach is carried out across an array of real and synthetic datasets that consistently demonstrate its exceptional performance and competitiveness ranking alongside state-of-the-art model-level explainers. Our code is available at https://github.com/amisayan/Gen-GraphEx. Sayan Saha 0001, Monidipa Das, Sanghamitra Bandyopadhyay |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Two-Dimensional Pheromone in Ant Colony OptimizationabstractAnt Colony Optimization (ACO) is an acclaimed method for solving combinatorial problems proposed by Marco Dorigo in 1992 and has since been enhanced and hybridized many times. This paper proposes a novel modification of the algorithm, based on the introduction of a two-dimensional pheromone into a single-criteria ACO. The complex structure of the pheromone is supposed to increase ants’ awareness when choosing the next edge of the graph, helping them achieve better results than in the original algorithm. The proposed modification is general and thus can be applied to any ACO-type algorithm. We show the results based on a representative instance of TSPLIB and discuss them in order to support our claims regarding the efficiency and efficacy of the proposed approach. Grazyna Starzec, Mateusz Starzec, Sanghamitra Bandyopadhyay, Ujjwal Maulik, Leszek Rutkowski, Marek Kisiel-Dorohinicki, Aleksander Byrski |
ICCCI | 3 |
| 2023 | IV-GNN : interval valued data handling using graph neural network
Sucheta Dawn, Sanghamitra Bandyopadhyay |
Appl. Intell. | 2 |
| 2023 | A scalable unsupervised learning of scRNAseq data detects rare cells through integration of structure-preserving embedding, clustering and outlier detectionabstractSingle-cell RNA-seq analysis has become a powerful tool to analyse the transcriptomes of individual cells. In turn, it has fostered the possibility of screening thousands of single cells in parallel. Thus, contrary to the traditional bulk measurements that only paint a macroscopic picture, gene measurements at the cell level aid researchers in studying different tissues and organs at various stages. However, accurate clustering methods for such high-dimensional data remain exiguous and a persistent challenge in this domain. Of late, several methods and techniques have been promulgated to address this issue. In this article, we propose a novel framework for clustering large-scale single-cell data and subsequently identifying the rare-cell sub-populations. To handle such sparse, high-dimensional data, we leverage PaCMAP (Pairwise Controlled Manifold Approximation), a feature extraction algorithm that preserves both the local and the global structures of the data and Gaussian Mixture Model to cluster single-cell data. Subsequently, we exploit Edited Nearest Neighbours sampling and Isolation Forest/One-class Support Vector Machine to identify rare-cell sub-populations. The performance of the proposed method is validated using the publicly available datasets with varying degrees of cell types and rare-cell sub-populations. On several benchmark datasets, the proposed method outperforms the existing state-of-the-art methods. The proposed method successfully identifies cell types that constitute populations ranging from 0.1 to 8% with F1-scores of 0.91 0.09. The source code is available at https://github.com/scrab017/RarPG. Koushik Mallick, Sikim Chakraborty, Saurav Mallik, Sanghamitra Bandyopadhyay |
Briefings Bioinform. | 4 |
| 2023 | SoURA: a user-reliability-aware social recommendation system based on graph neural network
Sucheta Dawn, Monidipa Das, Sanghamitra Bandyopadhyay |
Neural Comput. Appl. | 3 |
| 2022 | A Model-Centric Explainer for Graph Neural Network based Node ClassificationabstractGraph Neural Networks (GNNs) learn node representations by aggregating a node's feature vector with its neighbors. They perform well across a variety of graph tasks. However, to enhance the reliability and trustworthiness of these models during use in critical scenarios, it is of essence to look into the decision making mechanisms of these models rather than treating them as black boxes. Our model-centric method gives insight into the kind of information learnt by GNNs about node neighborhoods during the task of node classification. We propose a neighborhood generator as an explainer that generates optimal neighborhoods to maximize a particular class prediction of the trained GNN model. We formulate neighborhood generation as a reinforcement learning problem and use a policy gradient method to train our generator using feedback from the trained GNN-based node classifier. Our method provides intelligible explanations of learning mechanisms of GNN models on synthetic as well as real-world datasets and even highlights certain shortcomings of these models. Sayan Saha 0001, Monidipa Das, Sanghamitra Bandyopadhyay |
CIKM | 3 |
| 2022 | TURBaN: A Theory-Guided Model for Unemployment Rate Prediction Using Bayesian Network in Pandemic Scenario
Monidipa Das, Aysha Basheer, Sanghamitra Bandyopadhyay |
HIS | 3 |
| 2022 | Deep variational graph autoencoders for novel host-directed therapy options against COVID-19
Sumanta Ray, Snehalika Lall, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Alexander Schönhuth |
Artif. Intell. Medicine | 4 |
| 2022 | sc-REnF: An entropy guided robust feature selection for single-cell RNA-seq dataabstractAnnotation of cells in single-cell clustering requires a homogeneous grouping of cell populations. Since single-cell data are susceptible to technical noise, the quality of genes selected prior to clustering is of crucial importance in the preliminary steps of downstream analysis. Therefore, interest in robust gene selection has gained considerable attention in recent years. We introduce sc-REnF [robust entropy based feature (gene) selection method], aiming to leverage the advantages of $R{\prime}{e}nyi$ and $Tsallis$ entropies in gene selection for single cell clustering. Experiments demonstrate that with tuned parameter ($q$), $R{\prime}{e}nyi$ and $Tsallis$ entropies select genes that improved the clustering results significantly, over the other competing methods. sc-REnF can capture relevancy and redundancy among the features of noisy data extremely well due to its robust objective function. Moreover, the selected features/genes can able to determine the unknown cells with a high accuracy. Finally, sc-REnF yields good clustering performance in small sample, large feature scRNA-seq data. Availability: The sc-REnF is available at https://github.com/Snehalikalall/sc-REnF. Snehalika Lall, Abhik Ghosh, Sumanta Ray, Sanghamitra Bandyopadhyay |
Briefings Bioinform. | 4 |
| 2022 | GraMMy: Graph representation learning based on micro-macro analysis
Sucheta Dawn, Monidipa Das, Sanghamitra Bandyopadhyay |
Neurocomputing | 3 |
| 2022 | A copula based topology preserving graph convolution network for clustering of single-cell RNA-seq dataabstractAnnotation of cells in single-cell clustering requires a homogeneous grouping of cell populations. There are various issues in single cell sequencing that effect homogeneous grouping (clustering) of cells, such as small amount of starting RNA, limited per-cell sequenced reads, cell-to-cell variability due to cell-cycle, cellular morphology, and variable reagent concentrations. Moreover, single cell data is susceptible to technical noise, which affects the quality of genes (or features) selected/extracted prior to clustering. Here we introduce sc-CGconv (copula based graph convolution network for single clustering), a stepwise robust unsupervised feature extraction and clustering approach that formulates and aggregates cell-cell relationships using copula correlation (Ccor), followed by a graph convolution network based clustering approach. sc-CGconv formulates a cell-cell graph using Ccor that is learned by a graph-based artificial intelligence model, graph convolution network. The learned representation (low dimensional embedding) is utilized for cell clustering. sc-CGconv features the following advantages. a. sc-CGconv works with substantially smaller sample sizes to identify homogeneous clusters. b. sc-CGconv can model the expression co-variability of a large number of genes, thereby outperforming state-of-the-art gene selection/extraction methods for clustering. c. sc-CGconv preserves the cell-to-cell variability within the selected gene set by constructing a cell-cell graph through copula correlation measure. d. sc-CGconv provides a topology-preserving embedding of cells in low dimensional space. Snehalika Lall, Sumanta Ray, Sanghamitra Bandyopadhyay |
PLoS Comput. Biol. | 3 |
| 2022 | A Novel Graph Topology-Based GO-Similarity Measure for Signature Detection From Multi-Omics Data and its Application to Other ProblemsabstractLarge scale multi-omics data analysis and signature prediction have been a topic of interest in the last two decades. While various traditional clustering/correlation-based methods have been proposed, but the overall prediction is not always satisfactory. To solve these challenges, in this article, we propose a new approach by leveraging the Gene Ontology (GO)similarity combined with multiomics data. In this article, a new GO similarity measure, ModSchlicker, is proposed and the effectiveness of the proposed measure along with other standardized measures are reviewed while using various graph topology-based Information Content (IC)values of GO-term. The proposed measure is deployed to PPI prediction. Furthermore, by involving GO similarity, we propose a new framework for stronger disease-based gene signature detection from the multi-omics data. For the first objective, we predict interaction from various benchmark PPI datasets of Yeast and Human species. For the latter, the gene expression and methylation profiles are used to identify Differentially Expressed and Methylated (DEM)genes. Thereafter, the GO similarity score along with a statistical method are used to determine the potential gene signature. Interestingly, the proposed method produces a better performance ( 0.9 avg. accuracy and 0.95 AUC)as compared to the other existing related methods during the classification of the participating features (genes)of the signature. Moreover, the proposed method is highly useful in other prediction/classification problems for any kind of large scale omics data. Koushik Mallick, Saurav Mallik, Sanghamitra Bandyopadhyay, Sikim Chakraborty |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | A Multilayered Adaptive Recurrent Incremental Network Model for Heterogeneity-Aware Prediction of Derived Remote Sensing Image Time SeriesabstractCatastrophic forgettingof previously acquired knowledge is a major setback suffered by the neural networks (NNs) when these are trained on tasks in sequential fashion. The NN variants, including the deep network models, as commonly used in remote sensing data prediction are also not free from this limitation. The issue becomes more prominent when the prediction is performed over data collected from spatial zones with a large degree of subregional variations or heterogeneity. In order to tackle this problem, in this article, we propose a multilayered adaptive recurrent incremental network (MARINE) model, offering a spatial heterogeneity-aware self-adaptation scheme that resolves the catastrophic forgetting issue in an autonomous manner. Typically, the proposed MARINE is equipped with an intrinsic mechanism of clustering the spatial subregions as per their heterogeneity levels and auto-constructing the recurrent network layers in an ensemble fashion so that the knowledge acquired about one group of heterogeneity level does not overwrite that acquired about other groups. With respect to spatio-temporal prediction of normalized difference vegetation index (NDVI) time series, as derived from MODIS Terra satellite remote sensing imagery, we demonstrate that our proposed MARINE achieves competitive results when compared with the state-of-the-arts. More significantly, in the presence of higher degree of spatial heterogeneity, MARINE outperforms others by betteravoiding the catastrophic forgetting issue. Monidipa Das, Soumya K. Ghosh 0001, Sanghamitra Bandyopadhyay |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | S-conLSH: alignment-free gapped mapping of noisy long readsabstractBACKGROUND: The advancement of SMRT technology has unfolded new opportunities of genome analysis with its longer read length and low GC bias. Alignment of the reads to their appropriate positions in the respective reference genome is the first but costliest step of any analysis pipeline based on SMRT sequencing. However, the state-of-the-art aligners often fail to identify distant homologies due to lack of conserved regions, caused by frequent genetic duplication and recombination. Therefore, we developed a novel alignment-free method of sequence mapping that is fast and accurate. RESULTS: We present a new mapper called S-conLSH that uses Spaced context based Locality Sensitive Hashing. With multiple spaced patterns, S-conLSH facilitates a gapped mapping of noisy long reads to the corresponding target locations of a reference genome. We have examined the performance of the proposed method on 5 different real and simulated datasets. S-conLSH is at least 2 times faster than the recently developed method lordFAST. It achieves a sensitivity of 99%, without using any traditional base-to-base alignment, on human simulated sequence data. By default, S-conLSH provides an alignment-free mapping in PAF format. However, it has an option of generating aligned output as SAM-file, if it is required for any downstream processing. CONCLUSIONS: S-conLSH is one of the first alignment-free reference genome mapping tools achieving a high level of sensitivity. The spaced-context is especially suitable for extracting distant similarities. The variable-length spaced-seeds or patterns add flexibility to the proposed algorithm by introducing gapped mapping of the noisy long reads. Therefore, S-conLSH may be considered as a prominent direction towards alignment-free sequence analysis. Angana Chakraborty, Burkhard Morgenstern, Sanghamitra Bandyopadhyay |
BMC Bioinform. | 3 |
| 2021 | Novel weighted ensemble classifier for smartphone based indoor localization
Priya Roy, Chandreyee Chowdhury, Mausam Kundu, Dip Ghosh, Sanghamitra Bandyopadhyay |
Expert Syst. Appl. | 5 |
| 2021 | Supervised feature selection using integration of densest subgraph finding with floating forward-backward search
Tapas Bhadra, Sanghamitra Bandyopadhyay |
Inf. Sci. | 2 |
| 2021 | A detailed human activity transition recognition framework for grossly labeled data from smartphone accelerometer
Jayita Saha, Chandreyee Chowdhury, Dip Ghosh, Sanghamitra Bandyopadhyay |
Multim. Tools Appl. | 4 |
| 2021 | RgCop-A regularized copula based method for gene selection in single-cell RNA-seq dataabstractGene selection in unannotated large single cell RNA sequencing (scRNA-seq) data is important and crucial step in the preliminary step of downstream analysis. The existing approaches are primarily based on high variation (highly variable genes) or significant high expression (highly expressed genes) failed to provide stable and predictive feature set due to technical noise present in the data. Here, we propose RgCop, a novel regularized copula based method for gene selection from large single cell RNA-seq data. RgCop utilizes copula correlation (Ccor), a robust equitable dependence measure that captures multivariate dependency among a set of genes in single cell expression data. We formulate an objective function by adding l1 regularization term with Ccor to penalizes the redundant co-efficient of features/genes, resulting non-redundant effective features/genes set. Results show a significant improvement in the clustering/classification performance of real life scRNA-seq data over the other state-of-the-art. RgCop performs extremely well in capturing dependence among the features of noisy data due to the scale invariant property of copula, thereby improving the stability of the method. Moreover, the differentially expressed (DE) genes identified from the clusters of scRNA-seq data are found to provide an accurate annotation of cells. Finally, the features/genes obtained from RgCop is able to annotate the unknown cells with high accuracy. Snehalika Lall, Sumanta Ray, Sanghamitra Bandyopadhyay |
PLoS Comput. Biol. | 3 |
| 2021 | Stable feature selection using copula based mutual information
Snehalika Lall, Debajyoti Sinha, Abhik Ghosh, Debarka Sengupta, Sanghamitra Bandyopadhyay |
Pattern Recognit. | 5 |
| 2021 | Colored Network Motif Analysis by Dynamic Programming Approach: An Application in Host Pathogen Interaction NetworkabstractNetwork motifs are subgraphs of a network which are found with significantly higher frequency than that expected in similar random networks. Motifs are small building blocks of a network and they have emerged as a way to uncover topological properties of complex networks. A special yet not much explored type of motif is the 'colored motif' where color (type) of each node, and hence the edges, in the motif is distinguishable from each other. A traditional motif is defined as a recurring structure in a network, whereas colored motif introduces detailed information about the color of the nodes. G-trie is a data structure to efficiently store a given set of subgraphs by exploiting the topological overlaps within them. In this article we have implemented a modified g-trie to store colored subgraphs and developed a method to discover colored motifs. Our method uses an approximate enumeration for counting the subgraphs to reduce the runtime. We have applied our method to find colored motifs of size three in a host pathogen protein-protein interaction network having two types of proteins namely HIV-1 and human proteins, and four types of edges. Here, we have discovered eight motifs, six of which contain both HIV-1 and human proteins, while the remaining two contain only human proteins. Sumanta Ray, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Uniform distribution driven adaptive differential evolution
Raunak Sengupta, Monalisa Pal, Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Appl. Intell. | 4 |
| 2020 | Improved dropClust R package with integrative analysis support for scRNA-seq dataabstractSUMMARY: DropClust leverages Locality Sensitive Hashing (LSH) to speed up clustering of large scale single cell expression data. Here we present the improved dropClust, a complete R package that is, fast, interoperable and minimally resource intensive. The new dropClust features a novel batch effect removal algorithm that allows integrative analysis of single cell RNA-seq (scRNA-seq) datasets. AVAILABILITY AND IMPLEMENTATION: dropClust is freely available at https://github.com/debsin/dropClust as an R package. A lightweight online version of the dropClust is available at https://debsinha.shinyapps.io/dropClust/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Debajyoti Sinha, Pradyumn Sinha, Ritwik Saha, Sanghamitra Bandyopadhyay, Debarka Sengupta |
Bioinform. | 4 |
| 2020 | WeCoMXP: Weighted Connectivity Measure Integrating Co-Methylation, Co-Expression and Protein-Protein Interactions for Gene-Module DetectionabstractThe identification of modules (groups of several tightly interconnected genes) in gene interaction network is an essential task for better understanding of the architecture of the whole network. In this article, we develop a novel weighted connectivity measure integrating co-methylation, co-expression, and protein-protein interactions (called WeCoMXP) to detect gene-modules for multi-omics dataset. The proposed measure goes beyond the fundamental degree centrality measure through considering some formulation of higher-order connections. Thereafter, we apply the average linkage clustering method using the corresponding dissimilarity (distance) values of WeCoMXP scores, and utilize a dynamic tree cut method for identifying some gene-modules. We validate the modules through literature search, KEGG pathway, and gene-ontology analyses on the genes representing the modules. Furthermore, the top 10 TFs/miRNAs that are connected with the maximum number of gene-modules and that regulate/target the maximum number of genes from these connected gene-modules, are identified. Moreover, our proposed method provides a better performance than the existing methods in terms of several cluster-validity indices in maximum times. Saurav Mallik, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Topo2Vec: A Novel Node Embedding Generation Based on Network Topology for Link PredictionabstractLink prediction of a scale-free network has become relevant for problems relating to social network analysis, recommendation system, and in the domain of bioinformatics. In recently proposed approaches, the sampling of nodes of a network is done by simulating random walk. The generated node samples are used to train a neural network to learn the contextual information through an embedding vector. This method has gained popularity as the embedding vector is produced from the primitive node adjacency information and has achieved an outstanding performance. In this article, a naive and scalable approach for generating the node samples based on the principle of goal-oriented greedy searching has been proposed. The generated node samples have low noisy structures, which can represent the relation of edges of the network in a better way, compared to the state-of-the-art methods. Consequently, better representation of feature embedding of nodes of the network is generated from the samples. The learned feature vectors are used for solving the link prediction problem using a pairwise kernel support vector machine (SVM) classifier, which is computationally and spatially expensive in nature. Therefore, as an alternative, we have chosen to deploy a random forest (RF) classifier with a new algebraic operation to obtain the symmetric pairwise feature representation of a node pair. We demonstrated the efficacy of the proposed Topo2vec by testing it against the state-of-the-art network context generation algorithms in several real-world networks. Altogether, the proposed Topo2vec algorithm is a new method for solving the link prediction problem with entirely different settings. Koushik Mallick, Sanghamitra Bandyopadhyay, Subhasis Chakraborty, Rounaq Choudhuri, Sayan Bose |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2019 | Identification of Multiview Gene Modules Using Mutual Information-Based Hypograph MiningabstractDetection of gene-modules is one of the fundamental tasks for the integral analysis of network architecture. In this paper, we propose a novel algorithm using an integrated approach comprising statistical method and normalized mutual information-based hypograph mining for discovering the multiview co-similarity gene modules contained in multiview datasets. For this purpose, we first identify the statistically significant genes corresponding to each data profile and subsequently obtain the union set consisting of all these statistically significant genes. For each data profile, we then propose a new similarity score called as integrated normalized mutual information to obtain the similarity scores across all possible pairs of genes belonging to the union set by employing the results of gene clustering obtained through applying normalized mutual information-based graph clustering on the corresponding data profile. Moreover, we propose a new information theoretic measure called as multiview normalized mutual information to integrate all the similarity scores of a given gene-pair obtained across all the data profiles. For the experiment, we utilize one of the recently used multiview dataset named TCGA acute myeloid leukemia dataset comprising five different categories of data profiles. Furthermore, the co-similarity strengths of all the multiview gene modules obtained using the proposed method (PM) are reported. Finally, we provide a comparative study between the proposed and other existing methods for demonstrating the superiority of the PM over others. Code is available in http://ieeexplore.ieee.org. Tapas Bhadra, Saurav Mallik, Sanghamitra Bandyopadhyay |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2018 | DECOR: Differential Evolution using Clustering based Objective Reduction for many-objective optimization
Monalisa Pal, Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Inf. Sci. | 3 |
| 2018 | Integrating Multiple Data Sources for Combinatorial Marker Discovery: A Study in Tumorigenesis
Sanghamitra Bandyopadhyay, Saurav Mallik |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Multi-size patch based collaborative representation for Palm Dorsa Vein Pattern recognition by enhanced ensemble learning with modified interactive artificial bee colony algorithm
Sandip Joardar, Amitava Chatterjee, Sanghamitra Bandyopadhyay, Ujjwal Maulik |
Eng. Appl. Artif. Intell. | 3 |
| 2017 | A New Feature Vector Based on Gene Ontology Terms for Protein-Protein Interaction PredictionabstractProtein-protein interaction (PPI) plays a key role in understanding cellular mechanisms in different organisms. Many supervised classifiers like Random Forest (RF) and Support Vector Machine (SVM) have been used for intra or inter-species interaction prediction. For improving the prediction performance, in this paper we propose a novel set of features to represent a protein pair using their annotated Gene Ontology (GO) terms, including their ancestors. In our approach, a protein pair is treated as a document (bag of words), where the terms annotating the two proteins represent the words. Feature value of each word is calculated using information content of the corresponding term multiplied by a coefficient, which represents the weight of that term inside a document (i.e., a protein pair). We have tested the performance of the classifier using the proposed feature on different well known data sets of different species like S. cerevisiae, H. Sapiens, E. Coli, and D. melanogaster. We compare it with the other GO based feature representation technique, and demonstrate its competitive performance. Sanghamitra Bandyopadhyay, Koushik Mallick |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | A Scoring Scheme for Online Feature Selection: Simulating Model Performance Without RetrainingabstractIncreasing the number of features increases the complexity of a model even if the additional feature does not improve its decision-making capacity. Irrelevant features may also cause overfitting and reduce interpretability of the concerned model. It is, therefore, important that the features are optimally selected before a model is built. In the case of online learning, new instances are periodically discovered, and the respective model is tactically retrained as required. Similarly, there are many real-life situations where hundreds of new features are discovered periodically, and the existing model needs to be retrained or tested for its performance improvement. Supervised selection of feature subset usually requires creation of multiple suboptimal models, thus incurring time-intensive computations. Unsupervised selections, although faster, largely rely on some subjective definition of feature relevance. In this paper, we introduce a score that accurately determines the importance of the features. The proposed score is appropriate for online feature selection scenarios for its low time complexity and ability to interpret performance improvement of the current model after the addition of a new feature, without invoking a retraining. Debarka Sengupta, Sanghamitra Bandyopadhyay, Debajyoti Sinha |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Clustering based online automatic objective reduction to aid many-objective optimizationabstractStarting from scalability to visualization, several challenges come into play when multi-objective optimization algorithms are applied to many-objective optimization problems. These challenges can be tackled using objective reduction algorithms. This work proposes a correlation based objective reduction algorithm. The set of points representing the Pareto-Front are clustered based on correlation distance. The constituent objectives belonging to the most compact cluster are pruned except the centroid. The reduction and search proceed while alternating between full set and reduced set. Number of clusters is gradually increased in order to zoom in the correlation structure of the objectives and the algorithm stops when the number of clusters equals to the number of objectives. This objective reduction leads to automatic detection of number of objectives, faster convergence as one or more number of objectives can be eliminated in reduction stages and improved performance as both global exploration and local exploitation of the search space are considered. The proposed approach of objective reduction is coupled with a differential evolution based multi-objective optimization technique to develop a solution framework for many-objective optimization problems. The efficacy of the proposed approach is tested on several benchmark functions viz. DTLZ1, DTLZ2, DTLZ3 and DTLZ4 for 10 and 20 objectives. Convergence metric and hypervolume indicator are used for comparing the performance of the proposed solution framework with other popular multi-objective optimization algorithms. Monalisa Pal, Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
CEC | 3 |
| 2016 | A NMF based approach for integrating multiple data sources to predict HIV-1-human PPIsabstractBACKGROUND: Predicting novel interactions between HIV-1 and human proteins contributes most promising area in HIV research. Prediction is generally guided by some classification and inference based methods using single biological source of information. RESULTS: In this article we have proposed a novel framework to predict protein-protein interactions (PPIs) between HIV-1 and human proteins by integrating multiple biological sources of information through non negative matrix factorization (NMF). For this purpose, the multiple data sets are converted to biological networks, which are then utilized to predict modules. These modules are subsequently combined into meta-modules by using NMF based clustering method. The integrated meta-modules are used to predict novel interactions between HIV-1 and human proteins. We have analyzed the significant GO terms and KEGG pathways in which the human proteins of the meta-modules participate. Moreover, the topological properties of human proteins involved in the meta modules are investigated. We have also performed statistical significance test to evaluate the predictions. CONCLUSIONS: Here, we propose a novel approach based on integration of different biological data sources, for predicting PPIs between HIV-1 and human proteins. Here, the integration is achieved through non negative matrix factorization (NMF) technique. Most of the predicted interactions are found to be well supported by the existing literature in PUBMED. Moreover, human proteins in the predicted set emerge as 'hubs' and 'bottlenecks' in the analysis. Low p-value in the significance test also suggests that the predictions are statistically significant. Sumanta Ray, Sanghamitra Bandyopadhyay |
BMC Bioinform. | 2 |
| 2016 | Use of line based symmetry for developing cluster validity indices
Sudipta Acharya, Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Soft Comput. | 3 |
| 2016 | Discovering Condition Specific Topological Pattern Changes in Coexpression Network: An Application to HIV-1 ProgressionabstractThe natural progression of HIV-1 begins with a short acute retroviral syndrome which typically transit to chronic and clinical latency stages and subsequently progresses to a symptomatic, life-threatening immunodeficiency disease known as AIDS. Microarray analysis based on gene coexpression is widely used to investigate the coregulation pattern of a group (or cluster) of genes in a specific phenotype. Moreover, an investigation on the topological patterns across multiple phenotypes can facilitate the understanding of stage specific infection pattern of HIV-1 virus. Here, we develop a novel framework to identify topological patterns of gene co-expression network and detect changes of modular structure across different stages of HIV progression. This is achieved by comparing the topological and intramodular properties of HIV infection modules. To capture the diversity in modular structure, some topological, correlation based, and eigengene based measures are utilized here. We have applied a rank aggregation scheme to rank all the modules to provide a good agreement between these measures. Some novel transcription factors like 'FOXO1', 'GATA3', 'GFI1', 'IRF1', 'IRF7', 'MAX', 'STAT1', 'STAT3', 'XBP1', and 'YY1' that merge from the modules show significant change in expression pattern over HIV progression stages. Moreover, we have performed an eigengene based analysis to reveal the perturbation in modular structure across three stages of HIV-1 progression. Sumanta Ray, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | Ties that matterabstractOn-line social networks mostly allow individuals to extend friend requests to all forms of possible connections including those related to official purposes, interests, family relations, friendships, and acquaintances. One requires to mine relevant connections in order to make reliable and meaningful interpretations following network analysis. Most networks lack weight assignments that mark the strength of a connection. Thus there is a requirement of methods that can effectively identify essential edges from only the topological information available. The method discussed in this article identifies unique and high number of mutual connections through weighted self-information. The extracted skeleton network has highly reduced number of edges, still conserving the centrality distributions as far as possible. The method used is applied locally to each node, to extract connections relevant to every node. Results are demonstrated on five datasets which show that the proposed method is able to eliminate a large number of irrelevant edges. The method is also found to scale well to large datasets. Garisha Chowdhary, Sanghamitra Bandyopadhyay |
IEEE BigData | 2 |
| 2015 | A fuzzy citation-kNN algorithm for multiple instance learningabstractIn multiple instance learning (MIL) setting, instances are grouped together in different labeled bags and the classifier tries to learn the label of unknown bags or instances. This is significantly different from traditional supervised learning techniques where the instances are labeled itself. In this work, a fuzzy based citation-kNN technique, which uses modified Hausdorff distance between bags, is introduced. Introduction of a fuzzy distance measure helps to solve the problem of overlapping bags. Effect of false positive instances in a positive bag are also reduced by calculating a fuzzy class membership for the training bags. Experiments on drug discovery and image datasets show that the performance of the proposed algorithm (MI-FCKNN) is better than the traditional citation-kNN and competitive with most state-of-the-art algorithms. Dip Ghosh, Sanghamitra Bandyopadhyay |
FUZZ-IEEE | 2 |
| 2015 | A review of in silico approaches for analysis and prediction of HIV-1-human protein-protein interactionsabstractThe computational or in silico approaches for analysing the HIV-1-human protein-protein interaction (PPI) network, predicting different host cellular factors and PPIs and discovering several pathways are gaining popularity in the field of HIV research. Although there exist quite a few studies in this regard, no previous effort has been made to review these works in a comprehensive manner. Here we review the computational approaches that are devoted to the analysis and prediction of HIV-1-human PPIs. We have broadly categorized these studies into two fields: computational analysis of HIV-1-human PPI network and prediction of novel PPIs. We have also presented a comparative assessment of these studies and proposed some methodologies for discussing the implication of their results. We have also reviewed different computational techniques for predicting HIV-1-human PPIs and provided a comparative study of their applicability. We believe that our effort will provide helpful insights to the HIV research community. Sanghamitra Bandyopadhyay, Sumanta Ray, Anirban Mukhopadhyay 0001, Ujjwal Maulik |
Briefings Bioinform. | 1 |
| 2015 | Unsupervised feature selection using an improved version of Differential Evolution
Tapas Bhadra, Sanghamitra Bandyopadhyay |
Expert Syst. Appl. | 2 |
| 2015 | Priority based ∈ dominance: A new measure in multiobjective optimization
Sanghamitra Bandyopadhyay, Rudrasis Chakraborty, Ujjwal Maulik |
Inf. Sci. | 1 |
| 2015 | Finding quasi core with simulated stacked neural networks
Malay Bhattacharyya 0001, Sanghamitra Bandyopadhyay |
Inf. Sci. | 2 |
| 2015 | An Algorithm for Many-Objective Optimization With Reduced Objective Computations: A Study in Differential EvolutionabstractIn this paper we have developed an algorithm for many-objective optimization problems, which will work more quickly than existing ones, while offering competitive performance. The algorithm periodically reorders the objectives based on their conflict status and selects a subset of conflicting objectives for further processing. We have taken differential evolution multiobjective optimization (DEMO) as the underlying metaheuristic evolutionary algorithm, and implemented the technique of selecting a subset of conflicting objectives using a correlation-based ordering of objectives. The resultant method is called α-DEMO, where α is a parameter determining the number of conflicting objectives to be selected. We have also proposed a new form of elitism so as to restrict the number of higher ranked solutions that are selected in the next population. The α-DEMO with the revised elitism is referred to as α-DEMO-revised. Extensive results of the five DTLZ functions show that the number of objective computations required in the proposed algorithm is much less compared to the existing algorithms, while the convergence measures are competitive or often better. Statistical significance testing is also performed. A real-life application on structural optimization of factory shed truss is demonstrated. Sanghamitra Bandyopadhyay, Arpan Mukherjee |
IEEE Trans. Evol. Comput. | 1 |
| 2015 | Nonintrusive Load Monitoring: A Temporal Multilabel Classification ApproachabstractThe article tackles the issues related to the identification of electrical appliances inside residential buildings. Each appliance can be identified from the aggregate power readings at the meter panel. The possibility of applying a temporal multilabel classification approach in the domain of nonintrusive load monitoring is explored (nonevent-based method). A novel set of metafeatures is proposed. The method is tested on sampling rates based on the capabilities of current smart meters. The proposed approach is validated over a dataset of energy readings at residences for a period of a year for 100 houses containing different sets of appliances (water heater, washing machines, etc.). This method is applicable for the demand side management of households in the current limitation of smart meters; from the inhabitants or from the grid operator's point of view. Kaustav Basu, Vincent Debusschere, Seddik Bacha, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE Trans. Ind. Informatics | 5 |
| 2015 | FOCS: Fast Overlapped Community SearchabstractDiscovery of natural groups of similarly functioning individuals is a key task in analysis of real world networks. Also, overlap between community pairs is commonplace in large social and biological graphs, in particular. In fact, overlaps between communities are known to be denser than the non-overlapped regions of the communities. However, most of the existing algorithms that detect overlapping communities assume that the communities are denser than their surrounding regions, and falsely identify overlaps as communities. Further, many of these algorithms are computationally demanding and thus, do not scale reasonably with varying network sizes. In this article, we propose Fast Overlapped Community Search (FOCS), an algorithm that accounts for local connectedness in order to identify overlapped communities. FOCS is shown to be linear in number of edges and nodes. It additionally gains in speed via simultaneous selection of multiple near-best communities rather than merely the best, at each iteration. FOCS outperforms some popular overlapped community finding algorithms in terms of computational time while not compromising with quality. Sanghamitra Bandyopadhyay, Garisha Chowdhary, Debarka Sengupta |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Integration of dense subgraph finding with feature clustering for unsupervised feature selection
Sanghamitra Bandyopadhyay, Tapas Bhadra, Pabitra Mitra, Ujjwal Maulik |
Pattern Recognit. Lett. | 1 |
| 2014 | A Survey and Comparative Study of Statistical Tests for Identifying Differential Expressionfrom Microarray DataabstractDNA microarray is a powerful technology that can simultaneously determine the levels of thousands of transcripts (generated, for example, from genes/miRNAs) across different experimental conditions or tissue samples. The motto of differential expression analysis is to identify the transcripts whose expressions change significantly across different types of samples or experimental conditions. A number of statistical testing methods are available for this purpose. In this paper, we provide a comprehensive survey on different parametric and non-parametric testing methodologies for identifying differential expression from microarray data sets. The performances of the different testing methods have been compared based on some real-life miRNA and mRNA expression data sets. For validating the resulting differentially expressed miRNAs, the outcomes of each test are checked with the information available for miRNA in the standard miRNA database PhenomiR 2.0. Subsequently, we have prepared different simulated data sets of different sample sizes (from 10 to 100 per group/population) and thereafter the power of each test have been calculated individually. The comparative simulated study might lead to formulate robust and comprehensive judgements about the performance of each test in the basis of assumption of data distribution. Finally, a list of advantages and limitations of the different statistical tests has been provided, along with indications of some areas where further studies are required. Sanghamitra Bandyopadhyay, Saurav Mallik, Anirban Mukhopadhyay 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | A New Path Based Hybrid Measurefor Gene Ontology SimilarityabstractGene Ontology (GO) consists of a controlled vocabulary of terms, annotating a gene or gene product, structured in a directed acyclic graph. In the graph, semantic relations connect the terms, that represent the knowledge of functional description and cellular component information of gene products. GO similarity gives us a numerical representation of biological relationship between a gene set, which can be used to infer various biological facts such as protein interaction, structural similarity, gene clustering, etc. Here we introduce a new shortest path based hybrid measure of ontological similarity between two terms which combines both structure of the GO graph and information content of the terms. Here the similarity between two terms t1 and t2, referred to as GOSim(PBHM)(t1,t2), has two components; one obtained from the common ancestors of t1 and t2. The other from their remaining ancestors. The proposed path based hybrid measure does not suffer from the well-known shallow annotation problem. Its superiority with respect to some other popular measures is established for protein protein interaction prediction, correlation with gene expression and functional classification of genes in a biological pathway. Finally, the proposed measure is utilized to compute the average GO similarity score among the genes that are experimentally validated targets of some microRNAs. Results demonstrate that the targets of a given miRNA have a high degree of similarity in the biological process category of GO. Sanghamitra Bandyopadhyay, Koushik Mallick |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | A Survey of Multiobjective Evolutionary Algorithms for Data Mining: Part IabstractThe aim of any data mining technique is to build an efficient predictive or descriptive model of a large amount of data. Applications of evolutionary algorithms have been found to be particularly useful for automatic processing of large quantities of raw noisy data for optimal parameter setting and to discover significant and meaningful information. Many real-life data mining problems involve multiple conflicting measures of performance, or objectives, which need to be optimized simultaneously. Under this context, multiobjective evolutionary algorithms are gradually finding more and more applications in the domain of data mining since the beginning of the last decade. In this two-part paper, we have made a comprehensive survey on the recent developments of multiobjective evolutionary algorithms for data mining problems. In this paper, Part I, some basic concepts related to multiobjective optimization and data mining are provided. Subsequently, various multiobjective evolutionary approaches for two major data mining tasks, namely feature selection and classification, are surveyed. In Part II of this paper, we have surveyed different multiobjective evolutionary algorithms for clustering, association rule mining, and several other data mining tasks, and provided a general discussion on the scopes for future research in this domain. Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Carlos A. Coello Coello |
IEEE Trans. Evol. Comput. | 3 |
| 2014 | Survey of Multiobjective Evolutionary Algorithms for Data Mining: Part IIabstractThis paper is the second part of a two-part paper, which is a survey of multiobjective evolutionary algorithms for data mining problems. In Part I , multiobjective evolutionary algorithms used for feature selection and classification have been reviewed. In this part, different multiobjective evolutionary algorithms used for clustering, association rule mining, and other data mining tasks are surveyed. Moreover, a general discussion is provided along with scopes for future research in the domain of multiobjective evolutionary algorithms for data mining. Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Carlos A. Coello Coello |
IEEE Trans. Evol. Comput. | 3 |
| 2014 | Guest Editorial: Special Issue on Advances in Multiobjective Evolutionary Algorithms for Data MiningabstractThe six articles in this special issue provide a snapshot of the current research trends in multiobjective evolutionary algorithms for data mining. The main issues and challenges in this domain have been highlighted and directions of future research work have also been provided, Sanghamitra Bandyopadhyay, Ujjwal Maulik, Carlos A. Coello Coello, Witold Pedrycz |
IEEE Trans. Evol. Comput. | 1 |
| 2013 | Integrated analysis of gene expression and genome-wide DNA methylation for tumor prediction: An association rule mining-based approachabstractStatistical analysis and association rule mining are two most efficient techniques, where the first one is used to identify differentially expressed/methylated genes across different types of samples or experimental conditions and the second one is used to determine expression/methylation relationships among them. In this article, we have performed an integrated analysis of statistical methods and association rule mining on mRNA expression and DNA methylation datasets for the prediction of Uterine Leiomyoma. Moreover, we have proposed a novel rule-base classifier. Depending on 16 different rule-interestingness measures, we have applied a Genetic Algorithm based rank aggregation technique on the association rules which are generated from the training data by Apriori association rule mining algorithm. After determining the ranks of the rules, we have conducted a majority voting technique on each test point to determine its class-label (i.e. tumor or normal class-label) through weighted-sum method. We have run this classifier on the combined dataset using k-fold cross-validation and also performed a comparative performance analysis with other popular rule-base classifiers. Finally, we have predicted the status of some important genes (through frequency analysis in association rules for tumor and normal class-labels individually) that have a major role for tumor formation in Uterine Leiomyoma. Saurav Mallik, Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
CIBCB | 4 |
| 2013 | Incorporating fuzzy semantic similarity measure in detecting human protein complexes in PPI network: A multiobjective approachabstractDetection of protein complexes within protein-protein interaction networks (PPIN) is a valuable step toward the analysis of biological processes and pathways. Several high-throughput experimental techniques produce large number of PPIs that can be extensively utilized for constructing PPI network of a species. Decomposition of the whole PPI network into smaller and manageable modules is an ongoing challenge. Here we have developed a multi-objective algorithm for detecting human protein complexes by partitioning large human PPI network into clusters which serve as protein complexes. Some graphical properties like density, centrality etc., are utilized for building the objectives. Besides the graphical properties we have also exploited a fuzzy measure based semantic similarity approach to construct similarity based objective. The proposed technique is demonstrated in the human PPI network and the resulting complexes are analyzed in context of Gene Ontology (GO) and pathway enrichment. We have also compared our results with that of some state-of-the-art algorithms in context of different performance metrics. The biological relevance of our predicted complexes are also established here by linking them with 22 key disease classes. Sumanta Ray, Sanghamitra Bandyopadhyay, Anirban Mukhopadhyay 0001, Ujjwal Maulik |
FUZZ-IEEE | 2 |
| 2013 | Mining Quasi-Bicliques from HIV-1-Human Protein Interaction Network: A Multiobjective Biclustering ApproachabstractIn this work, we model the problem of mining quasi-bicliques from weighted viral-host protein-protein interaction network as a biclustering problem for identifying strong interaction modules. In this regard, a multiobjective genetic algorithm-based biclustering technique is proposed that simultaneously optimizes three objective functions to obtain dense biclusters having high mean interaction strengths. The performance of the proposed technique has been compared with that of other existing biclustering methods on an artificial data. Subsequently, the proposed biclustering method is applied on the records of biologically validated and predicted interactions between a set of HIV-1 proteins and a set of human proteins to identify strong interaction modules. For this, the entire interaction information is realized as a bipartite graph. We have further investigated the biological significance of the obtained biclusters. The human proteins involved in the strong interaction module have been found to share common biological properties and they are identified as the gateways of viral infection leading to various diseases. These human proteins can be potential drug targets for developing anti-HIV drugs. Ujjwal Maulik, Anirban Mukhopadhyay 0001, Malay Bhattacharyya 0001, Lars Kaderali, Benedikt Brors, Sanghamitra Bandyopadhyay, Roland Eils |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2013 | Reformulated Kemeny Optimal Aggregation with Application in Consensus Ranking of microRNA TargetsabstractMicroRNAs are very recently discovered small noncoding RNAs, responsible for negative regulation of gene expression. Members of this endogenous family of small RNA molecules have been found implicated in many genetic disorders. Each microRNA targets tens to hundreds of genes. Experimental validation of target genes is a time- and cost-intensive procedure. Therefore, prediction of microRNA targets is a very important problem in computational biology. Though, dozens of target prediction algorithms have been reported in the past decade, they disagree significantly in terms of target gene ranking (based on predicted scores). Rank aggregation is often used to combine multiple target orderings suggested by different algorithms. This technique has been used in diverse fields including social choice theory, meta search in web, and most recently, in bioinformatics. Kemeny optimal aggregation (KOA) is considered the more profound objective for rank aggregation. The consensus ordering obtained through Kemeny optimal aggregation incurs minimum pairwise disagreement with the input orderings. Because of its computational intractability, heuristics are often formulated to obtain a near optimal consensus ranking. Unlike its real time use in meta search, there are a number of scenarios in bioinformatics (e.g., combining microRNA target rankings, combining disease-related gene rankings obtained from microarray experiments) where evolutionary approaches can be afforded with the ambition of better optimization. We conjecture that an ideal consensus ordering should have its total disagreement shared, as equally as possible, with the input orderings. This is also important to refrain the evolutionary processes from getting stuck to local extremes. In the current work, we reformulate Kemeny optimal aggregation while introducing a trade-off between the total pairwise disagreement and its distribution. A simulated annealing-based implementation of the proposed objective has been found effective in context of microRNA target ranking. Supplementary data and source code link are available at: >http://www.isical.ac.in/bioinfo_miu/ieee_tcbb_kemeny.rar. Debarka Sengupta, Aroonalok Pyne, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2012 | Score Based Aggregation of microRNA Target Orderings
Debarka Sengupta, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
ISBRA | 3 |
| 2012 | Bounds on Quasi-Completeness
Malay Bhattacharyya 0001, Sanghamitra Bandyopadhyay |
IWOCA | 2 |
| 2012 | δ-TRIMAX: Extracting Triclusters and Analysing Coregulation in Time Series Gene Expression Data
Anirban Bhar, Martin Haubrock, Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Edgar Wingender |
WABI | 5 |
| 2012 | SVMeFC: SVM Ensemble Fuzzy Clustering for Satellite Image SegmentationabstractThe problem of unsupervised image segmentation of a satellite image in a number of homogeneous regions can be viewed as the task of clustering the pixels in the intensity space. This letter presents an approach that exploits the capability of some recently proposed fuzzy clustering techniques, as well as support vector machine (SVM) classifiers, to yield improved solutions. All the fuzzy clustering techniques are first used to produce a set of different clustering solutions. Each such solution has been improved by a novel technique based on an SVM classifier. Thereafter, the cluster-based similarity partition algorithm is used to create the final clustering solution from all improved ensemble solutions. Results demonstrating the effectiveness of the proposed technique are provided for numeric remote sensing data described in terms of feature vectors. Moreover, a remotely sensed image of Calcutta City has been segmented using the proposed technique to establish its utility. In addition, the additional information of this letter is given as supplementary at http://sysbio.icm.edu.pl/indra/SVMeFC.html. Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2012 | De Novo Design of Potential RecA Inhibitors Using MultiObjective OptimizationabstractDe novo ligand design involves optimization of several ligand properties such as binding affinity, ligand volume, drug likeness, etc. Therefore, optimization of these properties independently and simultaneously seems appropriate. In this paper, the ligand design problem is modeled in a multiobjective using Archived MultiObjective Simulated Annealing (AMOSA) as the underlying search algorithm. The multiple objectives considered are the energy components similarity to a known inhibitor and a novel drug likeliness measure based on Lipinski's rule of five. RecA protein of Mycobacterium tuberculosis, causative agent of tuberculosis, is taken as the target for the drug design. To gauge the goodness of the results, they are compared to the outputs of LigBuilder, NEWLEAD, and Variable genetic algorithm (VGA). The same problem has also been modeled using a well-established genetic algorithm-based multiobjective optimization technique, Nondominated Sorting Genetic Algorithm-II (NSGA-II), to find the efficacy of AMOSA through comparative analysis. Results demonstrate that while some small molecules designed by the proposed approach are remarkably similar to the known inhibitors of RecA, some new ones are discovered that may be potential candidates for novel lead molecules against tuberculosis. Soumi Sengupta, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Weighted Markov Chain Based Aggregation of Biomolecule OrderingsabstractThe scope and effectiveness of Rank Aggregation (RA) have already been established in contemporary bioinformatics research. Rank aggregation helps in meta-analysis of putative results collected from different analytic or experimental sources. For example, we often receive considerably differing ranked lists of genes or microRNAs from various target prediction algorithms or microarray studies. Sometimes combining them all, in some sense, yields more effective ordering of the set of objects. Also, assigning a certain level of confidence to each source of ranking is a natural demand of aggregation. Assignment of weights to the sources of orderings can be performed by experts. Several rank aggregation approaches like those based on Markov Chains (MCs), evolutionary algorithms, etc., exist in the literature. Markov chains, in general, are faster than the evolutionary approaches. Unlike the evolutionary computing approaches Markov chains have not been used for weighted aggregation scenarios. This is because of the absence of a formal framework of Weighted Markov Chain (WMC). In this paper, we propose the use of a modified version of MC4 (one of the Markov chains proposed by Dwork et al., 2001), followed by the weighted analog of local Kemenization for performing rank aggregation, where the sources of rankings can be prioritized by an expert. Effectiveness of the weighted Markov chain approach over the very recently proposed Genetic Algorithm (GA) and Cross-Entropy Monte Carlo (MC) algorithm-based techniques, has been established for gene orderings from microarray analysis and orderings of predicted microRNA targets. Debarka Sengupta, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2011 | Discovery of MicroRNA markers: An SVM-based multiobjective feature selection approachabstractMicroRNAs (miRNAs) are small non-coding RNAs that have been shown to play important roles in gene regulation and various biological processes. The abnormal expression of some specific miRNAs often results in the development of cancer. In this article, we have utilized a multiobjective genetic algorithm-based feature selection algorithm wrapped with support vector machine (SVM) classifier for selecting promising miRNAs having differential expression in benign and malignant tissue samples. Subsequently, the non-dominated sets of promising miRNAs are aggregated into a single most promising miRNA subset. Finally, the Signal-to-Noise Ratio (SNR) statistic has been applied on the obtained miRNA subset for identifying potential miRNA markers that distinguish the two classes (benign and malignant) of tissue samples. The performance has been demonstrated on four real-life miRNA expression datasets for different SVM kernel functions and the identified miRNA markers are reported. Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
CIBCB | 3 |
| 2011 | Automatic MR brain image segmentation using a multiseed based multiobjective clustering approach
Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Appl. Intell. | 2 |
| 2011 | Improvement of new automatic differential fuzzy clustering using SVM classifier for microarray analysis
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
Expert Syst. Appl. | 3 |
| 2011 | Use of a fuzzy granulation-degranulation criterion for assessing cluster validity
Sanghamitra Bandyopadhyay, Sriparna Saha 0001, Witold Pedrycz |
Fuzzy Sets Syst. | 1 |
| 2011 | Unsupervised and Supervised Learning Approaches Together for Microarray AnalysisabstractIn this article, a novel concept is introduced by using both unsupervised and supervised learning. For unsupervised learning, the problem of fuzzy clustering in microarray data as a multiobjective optimization is used, which simultaneously optimizes two internal fuzzy cluster validity indices to yield a set of Pareto-optimal clustering solutions. In this regards, a new multiobjective differential evolution based fuzzy clustering technique has been proposed. Subsequently, for supervised learning, a fuzzy majority voting scheme along with support vector machine is used to integrate the clustering information from all the solutions in the resultant Pareto-optimal set. The performances of the proposed clustering techniques have been demonstrated on five publicly available benchmark microarray data sets. A detail comparison has been carried out with multiobjective genetic algorithm based fuzzy clustering, multiobjective differential evolution based fuzzy clustering, single objective versions of differential evolution and genetic algorithm based fuzzy clustering as well as well known fuzzy c-means algorithm. While using support vector machine, comparative studies of the use of four different kernel functions are also reported. Statistical significance test has been done to establish the statistical superiority of the proposed multiobjective clustering approach. Finally, biological significance test has been carried out using a web based gene annotation tool to show that the proposed integrated technique is able to produce biologically relevant clusters of coexpressed genes. Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
Fundam. Informaticae | 3 |
| 2011 | A Biologically Inspired Measure for Coexpression AnalysisabstractTwo genes are said to be coexpressed if their expression levels have a similar spatial or temporal pattern. Ever since the profiling of gene microarrays has been in progress, computational modeling of coexpression has acquired a major focus. As a result, several similarity/distance measures have evolved over time to quantify coexpression similarity/dissimilarity between gene pairs. Of these, correlation coefficient has been established to be a suitable quantifier of pairwise coexpression. In general, correlation coefficient is good for symbolizing linear dependence, but not for nonlinear dependence. In spite of this drawback, it outperforms many other existing measures in modeling the dependency in biological data. In this paper, for the first time, we point out a significant weakness of the existing similarity/distance measures, including the standard correlation coefficient, in modeling pairwise coexpression of genes. A novel measure, called BioSim, which assumes values between -1 and +1 corresponding to negative and positive dependency and 0 for independency, is introduced. The computation of BioSim is based on the aggregation of stepwise relative angular deviation of the expression vectors considered. The proposed measure is analytically suitable for modeling coexpression as it accounts for the features of expression similarity, expression deviation and also the relative dependence. It is demonstrated how the proposed measure is better able to capture the degree of coexpression between a pair of genes as compared to several other existing ones. The efficacy of the measure is statistically analyzed by integrating it with several module-finding algorithms based on coexpression values and then applying it on synthetic and biological data. The annotation results of the coexpressed genes as obtained from gene ontology establish the significance of the introduced measure. By further extending the BioSim measure, it has been shown that one can effectively identify the variability in the expression patterns over multiple phenotypes. We have also extended BioSim to figure out pairwise differential expression pattern and coexpression dynamics. The significance of these studies is shown based on the analysis over several real-life data sets. The computation of the measure by focusing on stepwise time points also makes it effective to identify partially coexpressed genes. On the whole, we put forward a complete framework for coexpression analysis based on the BioSim measure. Sanghamitra Bandyopadhyay, Malay Bhattacharyya 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Multiobjective Simulated Annealing for Fuzzy Clustering With Stability and ValidityabstractThe present article deals with an application of multiobjective optimization (MOO) technique for fuzzy clustering. The problem of automatic clustering can be posed as one of the searching for a suitable set of cluster centers such that some measure of cluster validity is optimized. However, no single validity measure works equally well for different kinds of data sets. Moreover, it is extremely difficult to combine the different measures into one, since the modality for such combination may not be possible to ascertain, while sometimes the measures may be incompatible. Thus, it appears natural to keep a number of such measures separate, and to optimize them simultaneously by applying an MOO technique. Here, the mean value of a cluster validity index, computed over the partitionings obtained for different bootstrapped samples of the data, and a measure of its stability, that is defined newly in terms of the validity index, are used for optimization. The search capability of a recently proposed archived multiobjective simulated annealing algorithm called AMOSA is utilized for this purpose. The characteristic features of AMOSA are its concepts of amount of domination and archive in simulated annealing and situation-specific acceptance probabilities. The final solutions are stored in an archive, from which a single solution is chosen in a semisupervised manner that assumes the existence of some labeled points. Results are demonstrated for several artificial and real-life data sets. Comparison is made with another recent multiobjective evolutionary clustering scheme, in addition to a single objective approach and two classical methods. Sanghamitra Bandyopadhyay |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2010 | Simultaneous informative gene selection and clustering through multiobjective optimizationabstractClustering methods are used for unsupervised classification of tumor subclasses in microarray gene expression data sets organized in a fashion where the rows represent the tumor samples and columns represent the genes. Clustering algorithms can be very sensitive with respect to the set of features (genes) considered in the clustering process. It is important to select the set of informative and relevant genes to be used for clustering. In this article, a multiobjective genetic algorithm based technique has been proposed for performing the tasks of gene selection and fuzzy clustering simultaneously. A novel encoding technique is developed in this regard and the algorithm searches for the best cluster centers while minimizing the number of selected genes. The number of clusters is evolved automatically. The performance of the proposed technique has been illustrated on an artificial data set and compared with that of several other related feature selection/clustering approaches. Moreover its performance is demonstrated on two real life multi-class gene expression data sets viz., Brain tumor and Lung tumor data sets. Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Real-coded differential crisp clustering for MRI brain image segmentationabstractIn this paper, a segmentation technique of multi-spectral magnetic resonance image of the brain using a new differential evolution based crisp clustering is proposed. Real-coded encoding of the cluster centres is used for this purpose. Here assignments of points to different clusters are made based on the Euclidean distance. The proposed method is applied on several simulated T1-weighted, T2-weighted and proton density for normal and MS lesion magnetic resonance brain images. Superiority of the proposed method over genetic algorithm based crisp clustering, simulated annealing based crisp clustering, K-means and average linkage are demonstrated quantitatively. Segmentation obtained by differential evolution based crisp clustering technique is also compared with the available ground truth information. Also statistical analysis has been conducted to judge the effectiveness. Matlab version of the software is available at http://bio.icm.edu.pl/~darman/MRI. Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | PuTmiR: A database for extracting neighboring transcription factors of human microRNAsabstractBACKGROUND: Some of the recent investigations in systems biology have revealed the existence of a complex regulatory network between genes, microRNAs (miRNAs) and transcription factors (TFs). In this paper, we focus on TF to miRNA regulation and provide a novel interface for extracting the list of putative TFs for human miRNAs. A putative TF of an miRNA is considered here as those binding within the close genomic locality of that miRNA with respect to its starting or ending base pair on the chromosome. Recent studies suggest that these putative TFs are possible regulators of those miRNAs. DESCRIPTION: The interface is built around two datasets that consist of the exhaustive lists of putative TFs binding respectively in the 10 kb upstream region (USR) and downstream region (DSR) of human miRNAs. A web server, named as PuTmiR, is designed. It provides an option for extracting the putative TFs for human miRNAs, as per the requirement of a user, based on genomic locality, i.e., any upstream or downstream region of interest less than 10 kb. The degree distributions of the number of putative TFs and miRNAs against each other for the 10 kb USR and DSR are analyzed from the data and they explore some interesting results. We also report about the finding of a significant regulatory activity of the YY1 protein over a set of oncomiRNAs related to the colon cancer. CONCLUSION: The interface provided by the PuTmiR web server provides an important resource for analyzing the direct and indirect regulation of human miRNAs. While it is already an established fact that miRNAs are regulated by TFs binding to their USR, this database might possibly help to study whether an miRNA can also be regulated by the TFs binding to their DSR. Sanghamitra Bandyopadhyay, Malay Bhattacharyya 0001 |
BMC Bioinform. | 1 |
| 2010 | SFSSClass: an integrated approach for miRNA based tumor classificationabstractBACKGROUND: MicroRNA (miRNA) expression profiling data has recently been found to be particularly important in cancer research and can be used as a diagnostic and prognostic tool. Current approaches of tumor classification using miRNA expression data do not integrate the experimental knowledge available in the literature. A judicious integration of such knowledge with effective miRNA and sample selection through a biclustering approach could be an important step in improving the accuracy of tumor classification. RESULTS: In this article, a novel classification technique called SFSSClass is developed that judiciously integrates a biclustering technique SAMBA for simultaneous feature (miRNA) and sample (tissue) selection (SFSS), a cancer-miRNA network that we have developed by mining the literature of experimentally verified cancer-miRNA relationships and a classifier uncorrelated shrunken centroid (USC). SFSSClass is used for classifying multiple classes of tumors and cancer cell lines. In a part of the investigation, poorly differentiated tumors (PDT) having non diagnostic histological appearance are classified while training on more differentiated tumor (MDT) samples. The proposed method is found to outperform the best known accuracy in the literature on the experimental data sets. For example, while the best accuracy reported in the literature for classifying PDT samples is approximately 76.5%, the accuracy of SFSSClass is found to be approximately 82.3%. The advantage of incorporating biclustering integrated with the cancer-miRNA network is evident from the consistently better performance of SFSSClass (integration of SAMBA, cancer-miRNA network and USC) over USC (eg., approximately 70.5% for SFSSClass versus approximately 58.8% in classifying a set of 17 MDT samples from 9 tumor types, approximately 91.7% for SFSSClass versus approximately 75% in classifying 12 cell lines from 6 tumor types and approximately 82.3% for SFSSClass versus approximately 41.2% in classifying 17 PDT samples from 11 tumor types). CONCLUSION: In this article, we develop the SFSSClass algorithm which judiciously integrates a biclustering technique for simultaneous feature (miRNA) and sample (tissue) selection, the cancer-miRNA network and a classifier. The novel integration of experimental knowledge with computational tools efficiently selects relevant features that have high intra-class and low inter-class similarity. The performance of the SFSSClass is found to be significantly improved with respect to the other existing approaches. Ramkrishna Mitra, Sanghamitra Bandyopadhyay, Ujjwal Maulik, Michael Q. Zhang |
BMC Bioinform. | 2 |
| 2010 | A new multiobjective clustering technique based on the concepts of stability and symmetry
Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Knowl. Inf. Syst. | 2 |
| 2010 | Application of a Multiseed-Based Clustering Technique for Automatic Satellite Image SegmentationabstractThe problem of classifying an image into different homogeneous regions is viewed as a task of clustering the pixels in the intensity space. In this letter, a newly developed genetic clustering technique is used for automatically segmenting remote sensing satellite images. Each cluster is divided into several small hyperspherical subclusters, and the centers of all these small subclusters are encoded in a chromosome to represent the whole clustering. For assigning points to different clusters, these local subclusters are considered individually. For the purpose of objective function evaluation, these subclusters are merged appropriately to form a variable number of global clusters. A newly proposed point-symmetry-distance-based cluster validity index,Symindex, is used as a measure of the validity of the corresponding segment. The effectiveness of the proposed technique compared to a fuzzy C-means clustering technique, a recently proposed GAPS clustering withSym-index-based method, and a subtractive clustering technique is demonstrated in identifying different land cover regions from two numeric image data sets and a remote sensing image of a part of the city of Kolkata. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2010 | A symmetry based multiobjective clustering technique for automatic evolution of clusters
Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Pattern Recognit. | 2 |
| 2010 | Integrating Clustering and Supervised Learning for Categorical Data AnalysisabstractThe problem of fuzzy clustering of categorical data, where no natural ordering among the elements of a categorical attribute domain can be found, is an important problem in exploratory data analysis. As a result, a few clustering algorithms with focus on categorical data have been proposed. In this paper, a modified differential evolution (DE)-based fuzzy c-medoids (FCMdd) clustering of categorical data has been proposed. The algorithm combines both local as well as global information with adaptive weighting. The performance of the proposed method has been compared with those using genetic algorithm, simulated annealing, and the classical DE technique, besides the FCMdd, fuzzy k-modes, and average linkage hierarchical clustering algorithm for four artificial and four real life categorical data sets. Statistical test has been carried out to establish the statistical significance of the proposed method. To improve the result further, the clustering method is integrated with a support vector machine (SVM), a well-known technique for supervised learning. A fraction of the data points selected from different clusters based on their proximity to the respective medoids is used for training the SVM. The clustering assignments of the remaining points are thereafter determined using the trained classifier. The superiority of the integrated clustering and supervised learning approach has been demonstrated. Ujjwal Maulik, Sanghamitra Bandyopadhyay, Indrajit Saha |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2009 | Analysis of microarray data using multiobjective variable string length genetic fuzzy clusteringabstractIn this article, a novel multiobjective variable string length real coded genetic fuzzy clustering scheme for clustering microarray gene expression data has been proposed. The proposed technique automatically evolves the number of clusters along with the clustering result. The multiobjective variable string length clustering technique encodes the cluster centers in its chromosomes and simultaneously optimizes two fuzzy validity indices namely PBM index and Xie-Beni validity measure. In the final generation, it produces a set of non-dominated solutions, from which the best solution is selected using Silhouette index which is independent of the number of clusters. The corresponding chromosome length provides the number of clusters. The proposed method is applied on three publicly available real life gene expression data. Superiority of the proposed method over some other well known clustering algorithms has been demonstrated quantitatively. Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Ujjwal Maulik |
IEEE Congress on Evolutionary Computation | 2 |
| 2009 | Unsupervised cancer classification through SVM-boosted multiobjective fuzzy clustering with majority voting ensembleabstractIn this article, we have presented an unsupervised cancer classification technique based on multiobjective genetic fuzzy clustering of the tissue samples. In this regard, coordinate of the cluster centers have been encoded in the chromosomes and three fuzzy cluster validity indices are simultaneously optimized. Each solution of the resultant Pareto-optimal set has been boosted by a novel technique based on Support Vector Machine (SVM) classification. Finally, the clustering information possessed by the non-dominated solutions are combined through a majority voting ensemble technique to produce the final clustering solution. The performance of the proposed multiobjective clustering method has been compared to several other microarray clustering algorithms for three publicly available benchmark cancer data sets, viz., Leukemia, Colon cancer and Lymphoma data to establish its superiority. Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE Congress on Evolutionary Computation | 3 |
| 2009 | TargetMiner: microRNA target prediction with systematic identification of tissue-specific negative examplesabstractMOTIVATION: Prediction of microRNA (miRNA) target mRNAs using machine learning approaches is an important area of research. However, most of the methods suffer from either high false positive or false negative rates. One reason for this is the marked deficiency of negative examples or miRNA non-target pairs. Systematic identification of non-target mRNAs is still not addressed properly, and therefore, current machine learning approaches are compelled to rely on artificially generated negative examples for training. RESULTS: In this article, we have identified approximately 300 tissue-specific negative examples using a novel approach that involves expression profiling of both miRNAs and mRNAs, miRNA-mRNA structural interactions and seed-site conservation. The newly generated negative examples are validated with pSILAC dataset, which elucidate the fact that the identified non-targets are indeed non-targets.These high-throughput tissue-specific negative examples and a set of experimentally verified positive examples are then used to build a system called TargetMiner, a support vector machine (SVM)-based classifier. In addition to assessing the prediction accuracy on cross-validation experiments, TargetMiner has been validated with a completely independent experimental test dataset. Our method outperforms 10 existing target prediction algorithms and provides a good balance between sensitivity and specificity that is not reflected in the existing methods. We achieve a significantly higher sensitivity and specificity of 69% and 67.8% based on a pool of 90 feature set and 76.5% and 66.1% using a set of 30 selected feature set on the completely independent test dataset. In order to establish the effectiveness of the systematically generated negative examples, the SVM is trained using a different set of negative data generated using the method in Yousef et al. A significantly higher false positive rate (70.6%) is observed when tested on the independent set, while all other factors are kept the same. Again, when an existing method (NBmiRTar) is executed with the our proposed negative data, we observe an improvement in its performance. These clearly establish the effectiveness of the proposed approach of selecting the negative examples systematically. AVAILABILITY: TargetMiner is now available as an online tool at www.isical.ac.in/ approximately bioinfo_miu Sanghamitra Bandyopadhyay, Ramkrishna Mitra |
Bioinform. | 1 |
| 2009 | Analyzing miRNA co-expression networks to explore TF-miRNA regulationabstractBACKGROUND: Current microRNA (miRNA) research in progress has engendered rapid accumulation of expression data evolving from microarray experiments. Such experiments are generally performed over different tissues belonging to a specific species of metazoan. For disease diagnosis, microarray probes are also prepared with tissues taken from similar organs of different candidates of an organism. Expression data of miRNAs are frequently mapped to co-expression networks to study the functions of miRNAs, their regulation on genes and to explore the complex regulatory network that might exist between Transcription Factors (TFs), genes and miRNAs. These directions of research relating miRNAs are still not fully explored, and therefore, construction of reliable and compatible methods for mining miRNA co-expression networks has become an emerging area. This paper introduces a novel method for mining the miRNA co-expression networks in order to obtain co-expressed miRNAs under the hypothesis that these might be regulated by common TFs. RESULTS: Three co-expression networks, configured from one patient-specific, one tissue-specific and a stem cell-based miRNA expression data, are studied for analyzing the proposed methodology. A novel compactness measure is introduced. The results establish the statistical significance of the sets of miRNAs evolved and the efficacy of the self-pruning phase employed by the proposed method. All these datasets yield similar network patterns and produce coherent groups of miRNAs. The existence of common TFs, regulating these groups of miRNAs, is empirically tested. The results found are very promising. A novel visual validation method is also proposed that reflects the homogeneity as well as statistical properties of the grouped miRNAs. This visual validation method provides a promising and statistically significant graphical tool for expression analysis. CONCLUSION: A heuristic mining methodology that resembles a clustering motivation is proposed in this paper. However, there remains a basic difference between the mining method and a clustering approach. The heuristic approach can produce priority modules (PM) from an miRNA co-expression network, by employing a self-pruning phase, which are analyzed for statistical and biological significance. The mining algorithm minimizes the space/time complexity of the analysis, and also handles noise in the data. In addition, the mining method reveals promising results in the unsupervised analysis of TF-miRNA regulation. Sanghamitra Bandyopadhyay, Malay Bhattacharyya 0001 |
BMC Bioinform. | 1 |
| 2009 | Combining Pareto-optimal clusters using supervised learning for identifying co-expressed genesabstractBACKGROUND: The landscape of biological and biomedical research is being changed rapidly with the invention of microarrays which enables simultaneous view on the transcription levels of a huge number of genes across different experimental conditions or time points. Using microarray data sets, clustering algorithms have been actively utilized in order to identify groups of co-expressed genes. This article poses the problem of fuzzy clustering in microarray data as a multiobjective optimization problem which simultaneously optimizes two internal fuzzy cluster validity indices to yield a set of Pareto-optimal clustering solutions. Each of these clustering solutions possesses some amount of information regarding the clustering structure of the input data. Motivated by this fact, a novel fuzzy majority voting approach is proposed to combine the clustering information from all the solutions in the resultant Pareto-optimal set. This approach first identifies the genes which are assigned to some particular cluster with high membership degree by most of the Pareto-optimal solutions. Using this set of genes as the training set, the remaining genes are classified by a supervised learning algorithm. In this work, we have used a Support Vector Machine (SVM) classifier for this purpose. RESULTS: The performance of the proposed clustering technique has been demonstrated on five publicly available benchmark microarray data sets, viz., Yeast Sporulation, Yeast Cell Cycle, Arabidopsis Thaliana, Human Fibroblasts Serum and Rat Central Nervous System. Comparative studies of the use of different SVM kernels and several widely used microarray clustering techniques are reported. Moreover, statistical significance tests have been carried out to establish the statistical superiority of the proposed clustering approach. Finally, biological significance tests have been carried out using a web based gene annotation tool to show that the proposed method is able to produce biologically relevant clusters of co-expressed genes. CONCLUSION: The proposed clustering method has been shown to perform better than other well-known clustering algorithms in finding clusters of co-expressed genes efficiently. The clusters of genes produced by the proposed technique are also found to be biologically significant, i.e., consist of genes which belong to the same functional groups. This indicates that the proposed clustering method can be used efficiently to identify co-expressed genes in microarray gene expression data.Supplementary Website The pre-processed and normalized data sets, the matlab code and other related materials are available at http://anirbanmukhopadhyay.50webs.com/mogasvm.html. Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay |
BMC Bioinform. | 3 |
| 2009 | Mining the Largest Dense Vertexlet in a Weighted Scale-free GraphabstractAn important problem of knowledge discovery that has recently evolved in various reallife networks is identifying the largest set of vertices that are functionally associated. The topology of many real-life networks shows scale-freeness, where the vertices of the underlying graph follow a power-law degree distribution. Moreover, the graphs corresponding to most of the real-life networks are weighted in nature. In this article, the problem of finding the largest group or association of vertices that are dense (denoted as dense vertexlet) in a weighted scale-free graph is addressed. Density quantifies the degree of similarity within a group of vertices in a graph. The density of a vertexlet is defined in a novel way that ensures significant participation of all the vertices within the vertexlet. It is established that the problem is NP-complete in nature. An upper bound on the order of the largest dense vertexlet of a weighted graph, with respect to certain density threshold value, is also derived. Finally, an O(n $^2$ log n) (n denotes the number of vertices in the graph) heuristic graph mining algorithm that produces an approximate solution for the problem is presented. Sanghamitra Bandyopadhyay, Malay Bhattacharyya 0001 |
Fundam. Informaticae | 1 |
| 2009 | Some Symmetry Based ClassifiersabstractIn this paper, a novel point symmetry based pattern classifier (PSC) is proposed. A recently developed point symmetry based distance is utilized to determine the amount of point symmetry of a particular test pattern with respect to a class prototype. Kd-tree based nearest neighbor search is used for reducing the complexity of point symmetry distance computation. The proposed point symmetry based classifier is well-suited for classifying data sets having point symmetric classes, irrespective of any convexity, overlap or size. In order to classify data sets having line symmetry property, a line symmetry based classifier (LSC) along the lines of PSC is thereafter proposed in this paper. To measure the total amount of line symmetry of a particular point in a class, a new definition of line symmetry based distance is also provided. Proposed LSC preserves the advantages of PSC. The performance of PSC and LSC are demonstrated in classifying fourteen artificial and real-life data sets of varying complexities. For the purpose of comparison, k-NN classifier and the well-known support vector machine (SVM) based classifiers are executed on the data sets used here for the experiments. Statistical analysis, ANOVA, is also performed to compare the performance of these classification techniques. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Fundam. Informaticae | 2 |
| 2009 | Semi-GAPS: A Semi-supervised Clustering Method Using Point SymmetryabstractIn this paper, an evolutionary technique for the semi-supervised clustering is proposed. The proposed technique uses a point symmetry based distance measure. Semi-supervised classification uses aspects of both unsupervised and supervised learning to improve upon the performance of traditional classification methods. In this paper the existing point symmetry based genetic clustering technique, GAPS-clustering, is extended in two different ways to handle the semi-supervised classi- fication problem. The proposed semi-GAPS clustering algorithmis able to detect any type of clusters irrespective of shape, size and convexity as long as they possess the point symmetry property. Kdtree based nearest neighbor search is used to reduce the complexity of finding the closest symmetric point. Adaptive mutation and crossover probabilities are used. Experimental results demonstrate practical performance benefits of the methodology in detecting classes having symmetrical shapes in case of semi-supervised clustering. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Fundam. Informaticae | 2 |
| 2009 | MR Brain Image Segmentation Using A Multi-seed Based Automatic Clustering TechniqueabstractIn this paper, the automatic segmentation of multispectral magnetic resonance image of the brain is posed as a clustering problem in the intensity space. Thereafter an automatic clustering technique is proposed to solve this problem. The proposed real-coded variable string length genetic clustering technique (MCVGAPS clustering) is able to evolve the number of clusters present in the data set automatically. Each cluster is divided into several small hyperspherical subclusters and the centers of all these small sub-clusters are encoded in a string to represent the whole clustering. For assigning points to different clusters, these local sub-clusters are considered individually. For the purpose of objective function evaluation, these sub-clusters are merged appropriately to form a variable number of global clusters. A recently developed point symmetry distance based cluster validity index, Sym-index, is optimized to automatically evolve the appropriate number of clusters present in an MR brain image. The proposed method is applied on several simulated T1-weighted, T2- weighted and proton density normal and MS lesion magnetic resonance brain images. Superiority of the proposed method over Fuzzy C-means, Expectation Maximization clustering algorithms are demonstrated quantitatively. The automatic segmentation obtained by multiseed based multiobjective clustering technique (MCVGAPS) is also compared with the available ground truth information. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Fundam. Informaticae | 2 |
| 2009 | A new point symmetry based fuzzy genetic clustering technique for automatic evolution of clusters
Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Inf. Sci. | 2 |
| 2009 | A New Line Symmetry Distance and Its Application to Data Clustering
Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
J. Comput. Sci. Technol. | 2 |
| 2009 | A new multiobjective simulated annealing based clustering technique using symmetry
Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Pattern Recognit. Lett. | 2 |
| 2009 | Multiobjective Genetic Algorithm-Based Fuzzy Clustering of Categorical AttributesabstractRecently, the problem of clustering categorical data, where no natural ordering among the elements of a categorical attribute domain can be found, has been gaining significant attention from researchers. With the growing demand for categorical data clustering, a few clustering algorithms with focus on categorical data have recently been developed. However, most of these methods attempt to optimize a single measure of the clustering goodness. Often, such a single measure may not be appropriate for different kinds of datasets. Thus, consideration of multiple, often conflicting, objectives appears to be natural for this problem. Although we have previously addressed the problem of multiobjective fuzzy clustering for continuous data, these algorithms cannot be applied for categorical data where the cluster means are not defined. Motivated by this, in this paper a multiobjective genetic algorithm-based approach for fuzzy clustering of categorical data is proposed that encodes the cluster modes and simultaneously optimizes fuzzy compactness and fuzzy separation of the clusters. Moreover, a novel method for obtaining the final clustering solution from the set of resultant Pareto-optimal solutions in proposed. This is based on majority voting among Pareto front solutions followed byk-nn classification. The performance of the proposed fuzzy categorical data-clustering techniques has been compared with that of some other widely used algorithms, both quantitatively and qualitatively. For this purpose, various synthetic and real-life categorical datasets have been considered. Also, a statistical significance test has been conducted to establish the significant superiority of the proposed multiobjective approach. Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE Trans. Evol. Comput. | 3 |
| 2009 | Finding Multiple Coherent Biclusters in Microarray Data Using Variable String Length Multiobjective Genetic AlgorithmabstractMicroarray technology enables the simultaneous monitoring of the expression pattern of a huge number of genes across different experimental conditions. Biclustering in microarray data is an important technique that discovers a group of genes that are coregulated in a subset of conditions. Biclustering algorithms require to identify coherent and nontrivial biclusters, i.e., the biclusters should have low mean squared residue and high row variance. A multiobjective genetic biclustering technique is proposed here that optimizes these objectives simultaneously. A novel encoding scheme that uses variable chromosome length is developed. Moreover, a new quantitative measure to evaluate the goodness of the biclusters is proposed. The performance of the proposed algorithm has been evaluated on both simulated and real-life gene expression datasets, and compared with some other well-known biclustering techniques. Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2009 | Performance Evaluation of Some Symmetry-Based Cluster Validity IndexesabstractIdentification of the correct number of clusters is an important consideration in clustering where several cluster validity indexes, primarily utilizing the Euclidean distance, have been used in the literature. The property of symmetry is observed in most clustering solutions. In this paper, the symmetry versions of nine cluster validity indexes, namely, Davies-Bouldin index, Dunn index, generalized Dunn index, point symmetry (PS) index,Iindex, Xie-Beni index, FS index,Kindex, and SV index, are proposed. It is empirically established that incorporation of the property of symmetry significantly improves the capabilities of these indexes in identifying the appropriate number of clusters. A recently developed PS-based genetic clustering technique, GAPS clustering, is used as the underlying partitioning algorithm. Results on six artificially generated and five real-life datasets show that symmetry-distance-basedIindex performs the best as compared to all the other eight indexes. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2008 | Multiobjective fuzzy biclustering in microarray data: Method and a new performance measureabstractObjective of any biclustering algorithm in microarray data is to discover a subset of genes that are expressed similarly in a subset of conditions. The boundaries of biclusters usually overlap as genes and conditions may belong to different biclusters with different membership degrees. Hence the notion of fuzzy sets is useful for discovering such overlapping biclusters. In this article an attempt has been made to develop a multiobjective genetic algorithm based approach for probabilistic fuzzy biclustering that minimizes the residual and maximizes cluster size and expression profile variance. A novel variable string length encoding has been proposed in this regard that encodes multiple biclusters in a single string. Also a new performance measure that reflects how a bicluster is statistically distinguished from the background is proposed. Performance of the proposed algorithm has been compared with some well known biclustering algorithms. Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Michael Q. Zhang, Xuegong Zhang |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | Combining multiobjective fuzzy clustering and probabilistic ANN classifier for unsupervised pattern classification: Application to satellite image segmentationabstractAn important approach to unsupervised pixel classification in remote sensing satellite imagery is to use clustering in the spectral domain. In this article, a recently proposed multiobjective fuzzy clustering scheme has been combined with artificial neural networks (ANN) based probabilistic classifier to yield better performance. The multiobjective technique is first used to produce a set of non-dominated solutions. A part of these solutions having high confidence level are then used to train the ANN classifier. Finally the remaining solutions are classified using the trained classifier. The performance of this technique has been compared with that of some other well- known algorithms for two artificial data sets and a IRS satellite image of the city of Calcutta. Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Ujjwal Maulik |
IEEE Congress on Evolutionary Computation | 2 |
| 2008 | A Neuro-GA Approach for the Maximum Fuzzy Clique Problem
Sanghamitra Bandyopadhyay, Malay Bhattacharyya 0001 |
ICONIP (1) | 1 |
| 2008 | A New Principal Axis Based Line Symmetry Measurement and Its Application to Clustering
Sanghamitra Bandyopadhyay, Sriparna Saha 0001 |
ICONIP (2) | 1 |
| 2008 | A new multiobjective simulated annealing based clustering technique using stability and symmetryabstractMost clustering algorithms operate by optimizing (either implicitly or explicitly) a single measure of cluster solution quality. Such methods may perform well on some data sets but lack robustness with respect to variations in cluster shape, proximity, evenness and so forth. In this paper, we have proposed a multiobjective clustering technique which optimizes simultaneously two objectives, one reflecting the total symmetry present in the data set and the other reflecting the stability of the obtained partitions over different bootstrap samples of the data set. The proposed algorithm utilizes a recently developed simulated annealing based multiobjective optimization technique, AMOSA, as the underlying optimization method. Here assignment of points to different clusters are done based on the point symmetry based distance rather than the Euclidean distance. Results on several artificial and real-life data sets show that the proposed technique is well-suited to detect the number of clusters from data sets having point symmetric clusters. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
ICPR | 2 |
| 2008 | A multiobjective simulated annealing based fuzzy-clustering technique with symmetry for pixel classification in remote sensing imageryabstractAn important approach for landcover classification in remote sensing images is by clustering the pixels in the spectral domain into several fuzzy partitions. In this article the problem of fuzzy clustering is posed as one of searching for some suitable number of cluster centers so that some measures of validity of the obtained partitions should be optimized. In this paper a recently developed multiobjective simulated annealing based technique, AMOSA (archived multiobjective simulated annealing technique), is used to perform clustering, taking two validity measures as two objective functions. Here two fuzzy cluster validity functions namely, well-known XB-index and newly developed FSym-index are optimized simultaneously to automatically evolve the appropriate number of clusters present in an image. Thus the proposed algorithm provides a set of final non-dominated solutions, which the user can judge relatively and pick up the most promising one according to the domain requirement. The effectiveness of this proposed clustering technique in comparison with the existing fuzzy C-means clustering is shown for automatically classifying two remote sensing satellite images of the parts of the cities of Kolkata and Mumbai. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
ICPR | 2 |
| 2008 | A new line symmetry distance based pattern classifierabstractIn this paper, a new line symmetry based classifier (LSC) is proposed to deal with pattern classification problems. In order to measure total amount of line symmetry of a particular point in a class, a new definition of line symmetry based distance is also proposed in this paper. The proposed line symmetry based classifier (LSC) utilizes this new definition of line symmetry distance for classifying an unknown test sample. LSC assigns an unknown test sample pattern to that class with respect to whose major axis it is most symmetric. The mean of all the training patterns belonging to that particular class is taken as the prototype of that class. Thus training constitutes of computing only the class prototypes and the major axes of those classes. Kd-tree based nearest neighbor search is used for reducing complexity of line symmetry distance computation. The performance of LSC is demonstrated in classifying twelve artificial and real-life data sets of varying complexities. Experimental results show that LSC achieves, in general, higher classification accuracy compared to k-NN classifier. Results indicate that the proposed novel line symmetry based classifier is well-suited for classifying data sets having symmetrical classes, irrespective of any convexity, overlap and size. Statistical analysis, ANOVA is also performed to compare the performance of these classifications techniques. Sriparna Saha 0001, Sanghamitra Bandyopadhyay, Chingtham Tejbanta Singh |
IJCNN | 2 |
| 2008 | Fuzzy Symmetry Based Real-Coded Genetic Clustering Technique for Automatic Pixel Classification in Remote Sensing Imagery
Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
Fundam. Informaticae | 2 |
| 2008 | Application of a New Symmetry-Based Cluster Validity Index for Satellite Image SegmentationabstractAn important approach for image segmentation is clustering pixels based on their spectral properties. In particular, satellite images contain land cover types, some of which cover significantly large areas, while some (e.g., bridges and roads) occupy relatively much smaller regions. Automatically detecting regions or clusters of such widely varying sizes presents a challenging task. In this letter, a symmetry-based cluster validity index, named Sym-index (Symmetry distance-based index), is proposed. It is able to correctly indicate the presence of clusters of different sizes as long as they are internally symmetrical. A genetic-algorithm-based clustering technique that optimizes the Sym-index is used for image segmentation where the number of clusters is determined automatically. The superiority of the proposed index, as compared to other indices, is established for automatically segmenting the land cover types from SPOT and Indian Remote Sensing satellite images of two different cities in India. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2008 | A genetic approach for efficient outlier detection in projected space
Sanghamitra Bandyopadhyay, Santanu Santra |
Pattern Recognit. | 1 |
| 2008 | A Simulated Annealing-Based Multiobjective Optimization Algorithm: AMOSAabstractThis paper describes a simulated annealing based multiobjective optimization algorithm that incorporates the concept of archive in order to provide a set of tradeoff solutions for the problem under consideration. To determine the acceptance probability of a new solution vis-a-vis the current solution, an elaborate procedure is followed that takes into account the domination status of the new solution with the current solution, as well as those in the archive. A measure of the amount of domination between two solutions is also used for this purpose. A complexity analysis of the proposed algorithm is provided. An extensive comparative study of the proposed algorithm with two other existing and well-known multiobjective evolutionary algorithms (MOEAs) demonstrate the effectiveness of the former with respect to five existing performance measures, and several test problems of varying degrees of difficulty. In particular, the proposed algorithm is found to be significantly superior for many objective test problems (e.g., 4, 5, 10, and 15 objective problems), while recent studies have indicated that the Pareto ranking-based MOEAs perform poorly for such problems. In a part of the investigation, comparison of the real-coded version of the proposed algorithm is conducted with a very recent multiobjective simulated annealing algorithm, where the performance of the former is found to be generally superior to that of the latter. Sanghamitra Bandyopadhyay, Sriparna Saha 0001, Ujjwal Maulik, Kalyanmoy Deb |
IEEE Trans. Evol. Comput. | 1 |
| 2008 | A Point Symmetry-Based Clustering Technique for Automatic Evolution of ClustersabstractIn this article, a new symmetry based genetic clustering algorithm is proposed which automatically evolves the number of clusters as well as the proper partitioning from a data set. Strings comprise both real numbers and the don't care symbol in order to encode a variable number of clusters. Here, assignment of points to different clusters are done based on a point symmetry based distance rather than the Euclidean distance. A newly proposed point symmetry based cluster validity index, {\em Sym}-index, is used as a measure of the validity of the corresponding partitioning. The algorithm is therefore able to detect both convex and non-convex clusters irrespective of their sizes and shapes as long as they possess the point symmetry property. Kd-tree based nearest neighbor search is used to reduce the complexity of computing point symmetry based distance. A proof on the convergence property of variable string length GA with point symmetry based distance clustering (VGAPS-clustering) technique is also provided. The effectiveness of VGAPS-clustering compared to variable string length Genetic K-means algorithm (GCUK-clustering) and one recently developed weighted sum validity function based hybrid niching genetic algorithm (HNGA-clustering) is demonstrated for nine artificial and five real-life data sets. Sanghamitra Bandyopadhyay, Sriparna Saha 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | Gene Identification: Classical and Computational Intelligence ApproachesabstractAutomatic identification of genes has been an actively researched area of bioinformatics. Compared to earlier attempts for finding genes, the recent techniques are significantly more accurate and reliable. Many of the current gene-finding methods employ computational intelligence techniques that are known to be more robust when dealing with uncertainty and imprecision. In this paper, a detailed survey on the existing classical and computational intelligence based methods for gene identification is carried out. This includes a brief description of the classical and computational intelligence methods before discussing their applications to gene finding. In addition, a long list of available gene finders is compiled. For the convenience of the readers, the list is enhanced by mentioning their corresponding web sites and commenting on the general approach adopted. An extensive bibliography is provided. Finally, some limitations of the current approaches and future directions are discussed. Sanghamitra Bandyopadhyay, Ujjwal Maulik, Debadyuti Roy |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2007 | On some symmetry based validity indicesabstractIdentification of the correct number of clusters and the corresponding partitioning are two important considerations in clustering. In this paper, a newly developed point symmetry based distance is used to propose symmetry based versions of six cluster validity indices namely, DB-index, Dunn-index, generalized Dunn-index, PS-index, I-index and XB-index. These indices provide measures of "symmetricity" of the different partitionings of a data set. A Kd-tree-based data structure is used to reduce the complexity of computing the symmetry distance. A newly developed genetic point symmetry based clustering technique, GAPS-clustering is used as the underlying partitioning algorithm. The number of clusters are varied from 2 to radicn where n is the total number of data points present in the data set and the values of all the validity indices are noted down. The optimum value of a validity index over these radicn-1 partitions corresponds to the appropriate partitioning and the number of partitions as indicated by the validity index. Results on five artificially generated and four real-life data sets show that symmetry distance based I-index performs the best compared to all the other five indices. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
IEEE Congress on Evolutionary Computation | 2 |
| 2007 | MRI brain image segmentation by fuzzy symmetry based genetic clustering techniqueabstractIn this paper, an automatic segmentation technique of multispectral magnetic resonance image of the brain using a new fuzzy point symmetry based genetic clustering technique is proposed. The proposed real-coded variable string length genetic fuzzy clustering technique (fuzzy-VGAPS) is able to evolve the number of clusters present in the data set automatically. Here, assignment of points to different clusters are made based on the point symmetry based distance rather than the Euclidean distance. The cluster centers are encoded in the chromosomes, whose value may vary. A newly developed fuzzy point symmetry based cluster validity index, FSym-index, is used as a measure of 'goodness' of the corresponding partition. This validity index is able to correctly indicate presence of clusters of different sizes as long as they are internally symmetrical. A Kd-tree based data structure is used to reduce the complexity of computing the symmetry distance. The proposed method is applied on several simulated T1-weighted, T2-weighted and proton density normal and MS lesion magnetic resonance brain images. Superiority of the proposed method over fuzzy C-means, expectation maximization, fuzzy variable string length genetic algorithm (fuzzy-VGA) clustering algorithms are demonstrated quantitatively. The automatic segmentation obtained by fuzzy-VGAPS clustering technique is also compared with the available ground truth information. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
IEEE Congress on Evolutionary Computation | 2 |
| 2007 | Adaptive mufti-objective particle swarm optimization algorithmabstractIn this article we describe a novel Particle Swarm Optimization (PSO) approach to Multi-objective Optimization (MOO) called Adaptive Multi-objective Particle Swarm Optimization (AMOPSO). AMOPSO algorithm's novelty lies in its adaptive nature, that is attained by incorporating inertia and the acceleration coefficient as control variables with usual optimization variables, and evolving these through the swarming procedure. A new diversity parameter has been used to ensure sufficient diversity amongst the solutions of the non dominated front. AMOPSO has been compared with some recently developed multi-objective PSO techniques and evolutionary algorithms for nine function optimization problems, using different performance measures. Praveen Kumar Tripathi, Sanghamitra Bandyopadhyay, Sankar K. Pal |
IEEE Congress on Evolutionary Computation | 2 |
| 2007 | A validity index based on cluster symmetryabstractAn important consideration in clustering is the determination of the correct number of clusters and the appropriate partitioning of a given data set. In this paper, a newly developed point symmetry distance is used to propose a new cluster validity index named Sym -index which provides a measure of "symmetricity" of the different partitionings of a data set. The index is able to address all the above mentioned issues, viz., determining the number of clusters and evolving the proper partitioning as long as the clusters possess the property of symmetry. A Kd-tree-based data structure is used to reduce the complexity of computing the symmetry distance. Results demonstrating the superiority of the Sym-index in appropriately determining the proper partitioning and the number of clusters, as compared to two other recently proposed measures, namely the PS-index and I-index, are provided for three clustering methods viz., two recently developed genetic algorithm based clustering techniques and the average linkage clustering algorithm. Four artificial data sets and two real life data sets are considered for this purpose. The effectiveness of the proposed validity index is then demonstrated for automatically classifying different landcover regions in remote sensing imagery. Sriparna Saha 0001, Sanghamitra Bandyopadhyay |
SMC | 2 |
| 2007 | Genetic operators for combinatorial optimization in TSP and microarray gene ordering
Shubhra Sankar Ray, Sanghamitra Bandyopadhyay, Sankar K. Pal |
Appl. Intell. | 2 |
| 2007 | An improved algorithm for clustering gene expression dataabstractAbstract Motivation: Recent advancements in microarray technology allows simultaneous monitoring of the expression levels of a large number of genes over different time points. Clustering is an important tool for analyzing such microarray data, typical properties of which are its inherent uncertainty, noise and imprecision. In this article, a two-stage clustering algorithm, which employs a recently proposed variable string length genetic scheme and a multiobjective genetic clustering algorithm, is proposed. It is based on the novel concept of points having significant membership to multiple classes. An iterated version of the well-known Fuzzy C-Means is also utilized for clustering. Results: The significant superiority of the proposed two-stage clustering algorithm as compared to the average linkage method, Self Organizing Map (SOM) and a recently developed weighted Chinese restaurant-based clustering method (CRC), widely used methods for clustering gene expression data, is established on a variety of artificial and publicly available real life data sets. The biological relevance of the clustering solutions are also analyzed. Contact: [email protected] Supplementary information: The processed and normalized data sets, supplementary figures, tables and other related materials are available at http://d.1asphost.com/anirbanmukhopadhyay/simmts.html Sanghamitra Bandyopadhyay, Anirban Mukhopadhyay 0001, Ujjwal Maulik |
Bioinform. | 1 |
| 2007 | Multi-Objective Particle Swarm Optimization with time variant inertia and acceleration coefficients
Praveen Kumar Tripathi, Sanghamitra Bandyopadhyay, Sankar K. Pal |
Inf. Sci. | 2 |
| 2007 | GAPS: A clustering method using a new point symmetry-based distance measure
Sanghamitra Bandyopadhyay, Sriparna Saha 0001 |
Pattern Recognit. | 1 |
| 2007 | Multiobjective Genetic Clustering for Pixel Classification in Remote Sensing ImageryabstractAn important approach for unsupervised landcover classification in remote sensing images is the clustering of pixels in the spectral domain into several fuzzy partitions. In this paper, a multiobjective optimization algorithm is utilized to tackle the problem of fuzzy partitioning where a number of fuzzy cluster validity indexes are simultaneously optimized. The resultant set of near-Pareto-optimal solutions contains a number of nondominated solutions, which the user can judge relatively and pick up the most promising one according to the problem requirements. Real-coded encoding of the cluster centers is used for this purpose. Results demonstrating the effectiveness of the proposed technique are provided for numeric remote sensing data described in terms of feature vectors. Different landcover regions in remote sensing imagery have also been classified using the proposed technique to establish its efficiency Sanghamitra Bandyopadhyay, Ujjwal Maulik, Anirban Mukhopadhyay 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2007 | Dynamic Range-Based Distance Measure for Microarray Expressions and a Fast Gene-Ordering AlgorithmabstractThis investigation deals with a new distance measure for genes using their microarray expressions and a new algorithm for fast gene ordering without clustering. This distance measure is called "Maxrange distance," where the distance between two genes corresponding to a particular type of experiment is computed using a normalization factor, which is dependent on the dynamic range of the gene expression values of that experiment. The new gene-ordering method called "Minimal Neighbor" is based on the concept of nearest neighbor heuristic involving O(n2) time complexity. The superiority of this distance measure and the comparability of the ordering algorithm have been extensively established on widely studied microarray data sets by performing statistical tests. An interesting application of this ordering algorithm is also demonstrated for finding useful groups of genes within clusters obtained from a nonhierarchical clustering method like the self-organizing map. Shubhra Sankar Ray, Sanghamitra Bandyopadhyay, Sankar K. Pal |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | Clustering using Multi-objective Genetic Algorithm and its Application to Image SegmentationabstractThis article presents a multiobjective fuzzy genetic clustering technique employing real coded encoding of cluster centers. Recent research has shown that clustering techniques that optimize a single objective may not provide satisfactory result because no single validity measure works well on different kinds of data sets. This fact has motivated us to develop a multiobjective fuzzy genetic clustering method that optimizes multiple validity measures simultaneously. User can chose any partitioning result from the resultant set of non dominated solutions according to the problem requirements. A number of artificial and real-life data sets have been clustered using the proposed fuzzy clustering method. Also the proposed algorithm has been applied for segmentation of a remote sensing image to show its effectiveness in pixel classification. Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Ujjwal Maulik |
SMC | 2 |
| 2006 | Clustering distributed data streams in peer-to-peer environments
Sanghamitra Bandyopadhyay, Chris Giannella, Ujjwal Maulik, Hillol Kargupta, Kun Liu 0001, Souptik Datta |
Inf. Sci. | 1 |
| 2006 | Evolutionary computation in bioinformatics: a reviewabstractThis paper provides an overview of the application of evolutionary algorithms in certain bioinformatics tasks. Different tasks such as gene sequence analysis, gene mapping, deoxyribonucleic acid (DNA) fragment assembly, gene finding, microarray analysis, gene regulatory network analysis, phylogenetic trees, structure prediction and analysis of DNA, ribonucleic acid and protein, and molecular docking with ligand design are, first of all, described along with their basic features. The relevance of using evolutionary algorithms to these problems is then mentioned. These are followed by different approaches, along with their merits, for addressing some of the aforesaid tasks. Finally, some limitations of the current research activity are provided. An extensive bibliography is included Sankar K. Pal, Sanghamitra Bandyopadhyay, Shubhra Sankar Ray |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2005 | Automatic Determination of the Number of Fuzzy Clusters Using Simulated Annealing with Variable Representation
Sanghamitra Bandyopadhyay |
ISMIS | 1 |
| 2005 | An efficient technique for superfamily classification of amino acid sequences: feature extraction, fuzzy clustering and prototype selection
Sanghamitra Bandyopadhyay |
Fuzzy Sets Syst. | 1 |
| 2005 | A study of some fuzzy cluster validity indices, genetic clustering and application to pixel classification
Malay Kumar Pakhira, Sanghamitra Bandyopadhyay, Ujjwal Maulik |
Fuzzy Sets Syst. | 2 |
| 2005 | Simulated Annealing Using a Reversible Jump Markov Chain Monte Carlo Algorithm for Fuzzy ClusteringabstractIn this paper, an approach for automatically clustering a data set into a number of fuzzy partitions with a simulated annealing using a reversible jump Markov chain Monte Carlo algorithm is proposed. This is in contrast to the widely used fuzzy clustering scheme, the fuzzy c-means (FCM) algorithm, which requires the a priori knowledge of the number of clusters. The said approach performs the clustering by optimizing a cluster validity index, the Xie-Beni index. It makes use of the homogeneous reversible jump Markov chain Monte Carlo (RJMCMC) kernel as the proposal so that the algorithm is able to jump between different dimensions, i.e., number of clusters, until the correct value is obtained. Different moves, like birth, death, split, merge, and update, are used for sampling a candidate state given the current state. The effectiveness of the proposed technique in optimizing the Xie-Beni index and thereby determining the appropriate clustering is demonstrated for both artificial and real-life data sets. In a part of the investigation, the utility of the fuzzy clustering scheme for classifying pixels in an IRS satellite image of Kolkata is studied. A technique for reducing the computation efforts in the case of satellite image data is incorporated. Sanghamitra Bandyopadhyay |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | An automatic shape independent clustering technique
Sanghamitra Bandyopadhyay |
Pattern Recognit. | 1 |
| 2004 | Validity index for crisp and fuzzy clusters
Malay Kumar Pakhira, Sanghamitra Bandyopadhyay, Ujjwal Maulik |
Pattern Recognit. | 2 |
| 2004 | Multiobjective GAs, quantitative indices, and pattern classificationabstractThe concept of multiobjective optimization (MOO) has been integrated with variable length chromosomes for the development of a nonparametric genetic classifier which can overcome the problems, like overfitting/overlearning and ignoring smaller classes, as faced by single objective classifiers. The classifier can efficiently approximate any kind of linear and/or nonlinear class boundaries of a data set using an appropriate number of hyperplanes. While designing the classifier the aim is to simultaneously minimize the number of misclassified training points and the number of hyperplanes, and to maximize the product of class wise recognition scores. The concepts of validation set (in addition to training and test sets) and validation functional are introduced in the multiobjective classifier for selecting a solution from a set of nondominated solutions provided by the MOO algorithm. This genetic classifier incorporates elitism and some domain specific constraints in the search process, and is called the CEMOGA-Classifier (constrained elitist multiobjective genetic algorithm based classifier). Two new quantitative indices, namely, the purity and minimal spacing, are developed for evaluating the performance of different MOO techniques. These are used, along with classification accuracy, required number of hyperplanes and the computation time, to compare the CEMOGA-Classifier with other related ones. Sanghamitra Bandyopadhyay, Sankar K. Pal, B. Aruna |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | Efficient BIST design for sequential machines using FiF-FoF values in machine statesabstractThis paper introduces a novel BIST-quality metric termed as the FiF -- FoF (Fan-in-Factor & Fan-out-Factor) defined on FSM-states. Based on the FiF -- FoF analysis, an efficient scheme is presented that ensures all state codes appear with uniform likelyhood at the present state (PS) lines during the test phase. This results in higher fault efficiency in a BIST structure. Experimental results on MCNC benchmarks show that the scheme improves fault efficiency of sequential circuits significantly, with marginal area overhead. Samir Roy, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Subhadip Basu, Biplab K. Sikdar |
ASP-DAC | 3 |
| 2003 | Fuzzy partitioning using a real-coded variable-length genetic algorithm for pixel classificationabstractThe problem of classifying an image into different homogeneous regions is viewed as the task of clustering the pixels in the intensity space. Real-coded variable string length genetic fuzzy clustering with automatic evolution of clusters is used for this purpose. The cluster centers are encoded in the chromosomes, and the Xie-Beni index is used as a measure of the validity of the corresponding partition. The effectiveness of the proposed technique is demonstrated for classifying different landcover regions in remote sensing imagery. Results are compared with those obtained using the well-known fuzzy C-means algorithm. Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2002 | An evolutionary technique based on K-Means algorithm for optimal clustering in RN
Sanghamitra Bandyopadhyay, Ujjwal Maulik |
Inf. Sci. | 1 |
| 2002 | Performance Evaluation of Some Clustering Algorithms and Validity IndicesabstractIn this article, we evaluate the performance of three clustering algorithms, hard K-Means, single linkage, and a simulated annealing (SA) based technique, in conjunction with four cluster validity indices, namely Davies-Bouldin index, Dunn's index, Calinski-Harabasz index, and a recently developed index I. Based on a relation between the index I and the Dunn's index, a lower bound of the value of the former is theoretically estimated in order to get unique hard K-partition when the data set has distinct substructures. The effectiveness of the different validity indices and clustering methods in automatically evolving the appropriate number of clusters is demonstrated experimentally for both artificial and real-life data sets with the number of clusters varying from two to ten. Once the appropriate number of clusters is determined, the SA-based clustering technique is used for proper partitioning of the data into the said number of clusters. Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Genetic clustering for automatic evolution of clusters and application to image classification
Sanghamitra Bandyopadhyay, Ujjwal Maulik |
Pattern Recognit. | 1 |
| 2002 | Efficient prototype reordering in nearest neighbor classification
Sanghamitra Bandyopadhyay, Ujjwal Maulik |
Pattern Recognit. | 1 |
| 2001 | Clustering Using Simulated Annealing with Probabilistic RedistributionabstractAn efficient partitional clustering technique, called SAKM-clustering, that integrates the power of simulated annealing for obtaining minimum energy configuration, and the searching capability of K-means algorithm is proposed in this article. The clustering methodology is used to search for appropriate clusters in multidimensional feature space such that a similarity metric of the resulting clusters is optimized. Data points are redistributed among the clusters probabilistically, so that points that are farther away from the cluster center have higher probabilities of migrating to other clusters than those which are closer to it. The superiority of the SAKM-clustering algorithm over the widely used K-means algorithm is extensively demonstrated for artificial and real life data sets. Sanghamitra Bandyopadhyay, Ujjwal Maulik, Malay Kumar Pakhira |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2001 | SAFE: An Efficient Feature Extraction Technique
Ujjwal Maulik, Sanghamitra Bandyopadhyay, John C. Trinder |
Knowl. Inf. Syst. | 2 |
| 2001 | Pixel classification using variable string genetic algorithms with chromosome differentiationabstractThe concept of chromosome differentiation, commonly witnessed in nature as male and female sexes, is incorporated in genetic algorithms with variable length strings for designing a nonparametric classification methodology. Its significance in partitioning different landcover regions from satellite images, having complex/overlapping class boundaries, is demonstrated. The classifier is able to evolve automatically the appropriate number of hyperplanes efficiently for modeling any kind of class boundaries optimally. Merits of the system over the related ones are established through the use of several quantitative measure. Sanghamitra Bandyopadhyay, Sankar K. Pal |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2001 | Nonparametric genetic clustering: comparison of validity indicesabstractA variable-string-length genetic algorithm (GA) is used for developing a novel nonparametric clustering technique when the number of clusters is not fixed a-priori. Chromosomes in the same population may now have different lengths since they encode different number of clusters. The crossover operator is redefined to tackle the concept of variable string length. A cluster validity index is used as a measure of the fitness of a chromosome. The performance of several cluster validity indices, namely the Davies-Bouldin (1979) index, Dunn's (1973) index, two of its generalized versions and a recently developed index, in appropriately partitioning a data set, are compared. Sanghamitra Bandyopadhyay, Ujjwal Maulik |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2000 | Fault tolerant permutation mapping in multistage interconnection network
Ujjwal Maulik, Sanghamitra Bandyopadhyay, Siddhartha Bhattacharyya 0001 |
J. Syst. Archit. | 2 |
| 2000 | Genetic algorithm-based clustering technique
Ujjwal Maulik, Sanghamitra Bandyopadhyay |
Pattern Recognit. | 2 |
| 2000 | VGA-Classifier: design and applicationsabstractA method for pattern classification using genetic algorithms (GAs) has been recently described in Pal, Bandyopadhyay and Murthy (1998), where the class boundaries of a data set are approximated by a fixed number H of hyperplanes. As a consequence of fixing H a priori, the classifier suffered from the limitation of overfitting (or underfitting) the training data with an associated loss of its generalization capability. In this paper, we propose a scheme for evolving the value of H automatically using the concept of variable length strings/chromosomes. The crossover and mutation operators are newly defined in order to handle variable string lengths. The fitness function ensures primarily the minimization of the number of misclassified samples, and also the reduction of the number of hyperplanes. Based on an analogy between the classification principles of the genetic classifier and multilayer perceptron (with hard limiting neurons), a method for automatically determining the architecture and the connection weights of the latter is described. Sanghamitra Bandyopadhyay, Late C. A. Murthy, Sankar K. Pal |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1999 | Relation Between VGA-classifier and MLP : Determination of Network ArchitectureabstractAn analogy between a genetic algorithm based pattern classification scheme (where hyperplanes are used to approximate the class boundaries through searching) and multilayer perceptron (MLP) based classifier is established. Based on this, a method for Sanghamitra Bandyopadhyay, Sankar K. Pal |
Fundam. Informaticae | 1 |
| 1998 | Further Experimentations on the Scalability of the GEMGA
Hillol Kargupta, Sanghamitra Bandyopadhyay |
PPSN | 2 |
| 1998 | Incorporating Chromosome Differentaition in Genetic Algorithms
Sanghamitra Bandyopadhyay, Sankar K. Pal, Ujjwal Maulik |
Inf. Sci. | 1 |
| 1998 | Simulated Annealing Based Pattern Classification
Sanghamitra Bandyopadhyay, Sankar K. Pal, Late C. A. Murthy |
Inf. Sci. | 1 |
| 1998 | Pattern classification using genetic algorithms: Determination of H
Sanghamitra Bandyopadhyay, Late C. A. Murthy, Sankar K. Pal |
Pattern Recognit. Lett. | 1 |
| 1998 | Genetic algorithms for generation of class boundariesabstractA method is described for finding decision boundaries, approximated by piecewise linear segments, for classifying patterns in R(N),N>/=2, using an elitist model of genetic algorithms. It involves generation and placement of a set of hyperplanes (represented by strings) in the feature space that yields minimum misclassification. A scheme for the automatic deletion of redundant hyperplanes is also developed in case the algorithm starts with an initial conservative estimate of the number of hyperplanes required for modeling the decision boundary. The effectiveness of the classification methodology, along with the generalization ability of the decision boundary, is demonstrated for different parameter values on both artificial data and real life data sets having nonlinear/overlapping class boundaries. Results are compared extensively with those of the Bayes classifier, k-NN rule and multilayer perceptron. Sankar K. Pal, Sanghamitra Bandyopadhyay, Late C. A. Murthy |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1997 | Pattern classification with genetic algorithms: Incorporation of chromosome differentiation
Sanghamitra Bandyopadhyay, Sankar K. Pal |
Pattern Recognit. Lett. | 1 |
| 1996 | GA-based pattern classification: theoretical and experimental studiesabstractMerits of genetic algorithms (GAs), an efficient evolutionary searching paradigm, are utilized for pattern classification in /spl Rfr//sup N/ by fitting hyperplanes to model the decision boundaries in the feature space. Theoretical analysis establishes that as the size of the training set (n) goes towards infinity, the error probability and the decision boundary of the GA based classifier will approach those of Bayes (optimum) classifier. Sanghamitra Bandyopadhyay, Late C. A. Murthy, Sankar K. Pal |
ICPR | 1 |
| 1995 | Pattern classification with genetic algorithms
Sanghamitra Bandyopadhyay, Late C. A. Murthy, Sankar K. Pal |
Pattern Recognit. Lett. | 1 |