Ujjwal Maulik

dblp:93/3249 · DBLP profile ↗
← Back
105ranked-venue papers
19as first author
20since 2021 · last 2026
0000-0003-1167-0774ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorSystems, architecture and hardware · 4 · 1 first-author · 2 since 2021Theory of computation · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Attention-Enhanced Architecture with Multi-objective Hyperparameter Optimization for Efficient Lung Segmentation
Sudipta Kumar Biswas, Mrittika Chakraborty, Rangan Das, Ujjwal Maulik, Sanghamitra Bandyopadhyay
ACIIDS (2)4
2026 Mechanistic Traceability in Drug-Disease Association Discovery: A Deterministic Graph-RAG Approach
Arinjoy Pramanik, Sounak Ghosh, Rangan Das, Ujjwal Maulik
ACIIDS (1)4
2025 Indexing demand response potential at multiple temporal granularity using network theory based analysis
Kakuli Mishra, Srinka Basu, Ujjwal Maulik
Data Min. Knowl. Discov.3
2025 Lightweight deep learning models for aerial scene classification: A comprehensive survey
Suparna Dutta, Monidipa Das, Ujjwal Maulik
Eng. Appl. Artif. Intell.3
2025 A Lightweight Aerial Scene Classifier Based on Adaptive Fusion of MobileNet and CapsNet
abstract
Scene-level classification of aerial images is challenging due to varying object scales, positions, inter-class similarity, and intra-class diversity. Convolutional Neural Networks (CNNs) are effective for extracting useful feature maps, while Capsule Networks (CapsNets) are adept at recognizing the pose information of objects in an image. This work uses both models to ensure that aerial scenes are accurately classified. Additionally, considering that CapsNets typically require many parameters and floating-point operations, we aim to make our model lightweight. Using adaptive fusion of features extracted from MobileNetv2, a lightweight CNN model, enables the capsule network to learn faster, resulting in improved performance. With more than 58% parameter reduction for the datasets UC Merced, AID, and SIRI-WHU, the proposed approach, termed as AdMobiCaps, achieves significantly better performance compared to traditional CapsNet showcasing its potential in aerial scene classification.
Suparna Dutta, Monidipa Das, Ujjwal Maulik
IEEE Signal Process. Lett.3
2024 An entropy-based membership approach on type-II fuzzy set (EMT2FCM) for biomedical image segmentation
Ananya Bose, Ujjwal Maulik, Anasua Sarkar
Eng. Appl. Artif. Intell.2
2024 A Generalized Attention Mechanism to Enhance the Accuracy Performance of Neural Networks
abstract
In many modern machine learning (ML) models, attention mechanisms (AMs) play a crucial role in processing data and identifying significant parts of the inputs, whether these are text or images. This selective focus enables subsequent stages of the model to achieve improved classification performance. Traditionally, AMs are applied as a preprocessing substructure before a neural network, such as in encoder/decoder architectures. In this paper, we extend the application of AMs to intermediate stages of data propagation within ML models. Specifically, we propose a generalized attention mechanism (GAM), which can be integrated before each layer of a neural network for classification tasks. The proposed GAM allows for at each layer/step of the ML architecture identification of the most relevant sections of the intermediate results. Our experimental results demonstrate that incorporating the proposed GAM into various ML models consistently enhances the accuracy of these models. This improvement is achieved with only a marginal increase in the number of parameters, which does not significantly affect the training time.
Pengcheng Jiang, Ferrante Neri, Yu Xue 0003, Ujjwal Maulik
Int. J. Neural Syst.4
2024 Toward Causality-Based Explanation of Aerial Scene Classifiers
abstract
Recently, convolutional neural networks (CNNs) have achieved great success by attaining state-of-the-art accuracies for aerial scene classification. However, there is a serious lack of good explanations for understanding the decision-making process of such black-box models. To establish the trustworthiness of the classifiers, various local and global explainers are commonly used nowadays. These primarily provide us with explanations in terms of the most influential features leading to the model decision. However, developing an explainer showing causal relationships among these features can offer more visibility, interpretability, and trustworthiness to the model users or stakeholders. To the best of our knowledge, this area is still underexplored in the context of scene-level classification of aerial images. We address this issue by proposing a novel causality-based CNN explainer based on gradient-weighted class activation mapping (Grad-CAM) and generative flow network (GFlowNet). Our proposed model is termed causal grad-CAM (CG-CAM), where Grad-CAM is used to highlight the relevant regions (layer-wise) in an aerial scene, and the GFlowNet is utilized to generate the directed acyclic graph (DAG) representing causal relationships between the feature maps at different layers of a deep CNN classifier, to achieve better understandability. Experimentation using the benchmark UCMerced and NWPU-RESISC45 datasets demonstrates the effectiveness of our CG-CAM-based explanations for aerial scene classification.
Suparna Dutta, Monidipa Das, Ujjwal Maulik
IEEE Geosci. Remote. Sens. Lett.3
2024 Disastrous Event and Sub-Event Detection From Microblog Posts Using Bi-Clustering Method
abstract
Social media has become a nondetachable part of our life, with the exponential growth of usage in the past decade. Social sites like Twitter, Facebook, Instagram, Flickr, Weibo, etc., with their millions of user base, apart from being a source of entertainment, has proven to be a very useful mean for public opinion generation, news propagation and information broadcasting by authorities. Social media data analysis has been a popular research area for the past few years. Detecting subevents from social media posts to identify an unusual event that requires special attention, especially in a disaster situation, is one of the key researches in this domain. In this article, we have proposed a novel biclustering-based subevent detection method from the Twitter dataset for retrospective analysis of disaster events. First, we have clustered the data matrix using spectral co-clustering. Then we identified subevents (words) and formulated a ranking framework to find the top-ranked subevents within the clusters. Finally, through statistical analysis, we have shown that the proposed framework works better than other existing subevent detection methods.
Shatadru Roy Chowdhury, Srinka Basu, Ujjwal Maulik
IEEE Trans. Comput. Soc. Syst.3
2023 Two-Dimensional Pheromone in Ant Colony Optimization
abstract
Ant Colony Optimization (ACO) is an acclaimed method for solving combinatorial problems proposed by Marco Dorigo in 1992 and has since been enhanced and hybridized many times. This paper proposes a novel modification of the algorithm, based on the introduction of a two-dimensional pheromone into a single-criteria ACO. The complex structure of the pheromone is supposed to increase ants’ awareness when choosing the next edge of the graph, helping them achieve better results than in the original algorithm. The proposed modification is general and thus can be applied to any ACO-type algorithm. We show the results based on a representative instance of TSPLIB and discuss them in order to support our claims regarding the efficiency and efficacy of the proposed approach.
Grazyna Starzec, Mateusz Starzec, Sanghamitra Bandyopadhyay, Ujjwal Maulik, Leszek Rutkowski, Marek Kisiel-Dorohinicki, Aleksander Byrski
ICCCI4
2023 Automatic Vehicle Pollution Detection Using Feedback Based Iterative Deep Learning
abstract
Air pollution is one of the major health hazards in modern times. Vehicle pollution is one of the major contributors to aerial contamination. Significant emphasis has been given by the researchers to identify the traffic pollutant sources. However, there is a need for a low-cost, automated solution for fast detection of polluting vehicles from the end of law enforcers. In this article, we have presented a novel deep learning based evaluation strategy that will identify the pollutant vehicle through on- road installed surveillance camera images. An enhanced image data set with notable variations has been prepared for training the models. Thereafter, a multi-model feedback process has been integrated. In contrast to the other deep learning approaches, the proposed framework has started with low labeled training data which is quite relevant in real-life scenarios. Subsequently, the size of the training samples has been increased using feedback. The evaluation of the framework has been performed with seven popular deep learning CNN models, i.e Inception-V3, MobileNet-V2, MobileNet-V3 (small), InceptionResNet-V2, VGG16, VGG19 and XceptionNet. Transfer learning has been exploited to distinguish the on- road pollutants and to provide proper surveillance in transport system. We have compared our method with recent state-of-the-art techniques. The results have demonstrated the superiority of the proposed framework in the road surveillance domain. To the best of our knowledge, this is the first attempt to apply a feedback based iterative deep learning model for vehicle pollution detection.
Ujjwal Maulik, Srimanta Kundu
IEEE Trans. Intell. Transp. Syst.1
2023 Cuckoo search optimization-based energy efficient job scheduling approach for IoT-edge environment
Mohana Bakshi, Chandreyee Chowdhury, Ujjwal Maulik
J. Supercomput.3
2022 Study of transcription factor druggabilty for prostate cancer using structure information, gene regulatory networks and protein moonlighting
abstract
Prostate cancer is the second leading cause of cancer-related death in men. Metastasis shows poor survival even though the recovery rate is high. In spite of numerous studies regarding prostate carcinoma, multiple questions are still unanswered. In this regards, gene regulatory network can uncover the mechanisms behind cancer progression, and metastasis. Under a feed forward loop, transcription factors (TFs) can be a good druggable candidate. We have proposed a computational model to study the uncertainty of TFs and suggest the appropriate cellular conditions for drug targeting. We have selected feed-forward loops depending on the shared list of the functional annotations among TFs, genes and miRNAs. From the potential feed forward loop cores, six TFs were identified as druggable targets, which include AR, CEBPB, CREB1, ETS1, NFKB1 and RELA. However, TFs are known for their Protein Moonlighting properties, which provide unrelated multi-functionalities within the same or different subcellular localizations. Following that, we have identified such functions that are suitable for drug targeting. On the other hand, we have tried to identify membraneless organelles for providing more specificity to the proposed time and space theory. The study has provided certain possibilities on TF-based therapeutics. The controlled dynamic nature of the TF may have enhanced the chances where TFs can be considered as one of the prime drug targets. Finally, the combination of membranless phase separation and protein moonlighting has provided possible druggable period within the biological clock.
Ashmita Dey, Sagnik Sen 0002, Ujjwal Maulik
Briefings Bioinform.3
2022 Graft: A graph based time series data mining framework
Kakuli Mishra, Srinka Basu, Ujjwal Maulik
Eng. Appl. Artif. Intell.3
2021 Negatively-Associated Maximal Frequent Geneset Mining on DNA Methylation Profile
abstract
Association rule mining has been an important approach for feature and biomarker discovery in various omics data. One main challenge is that it generates a large number of itemsets. The effect of this shortcoming increases substantially in the case of negative association rule mining that is useful for detecting significant relationships between genes (items/features) in the form of either presence or absence in disease characterization. In this article, we propose a new algorithm, NegaMax (negatively-associated maximal frequent itemsets) for negative association itemset mining. Our method follows depth-first search rather than breadth-first search used in the other methods. It identifies a much fewer number of non-redundant itemsets than that by the existing methods. Thus, it saves elapsing time for itemset generation which potentially remove false positive intermediate results. We demonstrated NegaMax in a real-world DNA methylation dataset. The proposed method is highly beneficial from a medical perspective.
Saurav Mallik, Souvik Rakshit, Ujjwal Maulik, Zhongming Zhao
BIBM3
2021 Understanding structural malleability of the SARS-CoV-2 proteins and relation to the comorbidities
abstract
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), a causative agent of the coronavirus disease (COVID-19), is a part of the $\beta $-Coronaviridae family. The virus contains five major protein classes viz., four structural proteins [nucleocapsid (N), membrane (M), envelop (E) and spike glycoprotein (S)] and replicase polyproteins (R), synthesized as two polyproteins (ORF1a and ORF1ab). Due to the severity of the pandemic, most of the SARS-CoV-2-related research are focused on finding therapeutic solutions. However, studies on the sequences and structure space throughout the evolutionary time frame of viral proteins are limited. Besides, the structural malleability of viral proteins can be directly or indirectly associated with the dysfunctionality of the host cell proteins. This dysfunctionality may lead to comorbidities during the infection and may continue at the post-infection stage. In this regard, we conduct the evolutionary sequence-structure analysis of the viral proteins to evaluate their malleability. Subsequently, intrinsic disorder propensities of these viral proteins have been studied to confirm that the short intrinsically disordered regions play an important role in enhancing the likelihood of the host proteins interacting with the viral proteins. These interactions may result in molecular dysfunctionality, finally leading to different diseases. Based on the host cell proteins, the diseases are divided in two distinct classes: (i) proteins, directly associated with the set of diseases while showing similar activities, and (ii) cytokine storm-mediated pro-inflammation (e.g. acute respiratory distress syndrome, malignancies) and neuroinflammation (e.g. neurodegenerative and neuropsychiatric diseases). Finally, the study unveils that males and postmenopausal females can be more vulnerable to SARS-CoV-2 infection due to the androgen-mediated protein transmembrane serine protease 2.
Sagnik Sen 0002, Ashmita Dey, Sanghamitra Bandhyopadhyay, Vladimir N. Uversky, Ujjwal Maulik
Briefings Bioinform.5
2021 Unveiling COVID-19-associated organ-specific cell types and cell-specific pathway cascade
abstract
The novel coronavirus or COVID-19 has first been found in Wuhan, China, and became pandemic. Angiotensin-converting enzyme 2 (ACE2) plays a key role in the host cells as a receptor of Spike-I Glycoprotein of COVID-19 which causes final infection. ACE2 is highly expressed in the bladder, ileum, kidney and liver, comparing with ACE2 expression in the lung-specific pulmonary alveolar type II cells. In this study, the single-cell RNAseq data of the five tissues from different humans are curated and cell types with high expressions of ACE2 are identified. Subsequently, the protein-protein interaction networks have been established. From the network, potential biomarkers which can form functional hubs, are selected based on k-means network clustering. It is observed that angiotensin PPAR family proteins show important roles in the functional hubs. To understand the functions of the potential markers, corresponding pathways have been researched thoroughly through the pathway semantic networks. Subsequently, the pathways have been ranked according to their influence and dependency in the network using PageRank algorithm. The outcomes show some important facts in terms of infection. Firstly, renin-angiotensin system and PPAR signaling pathway can play a vital role for enhancing the infection after its intrusion through ACE2. Next, pathway networks consist of few basic metabolic and influential pathways, e.g. insulin resistance. This information corroborate the fact that diabetic patients are more vulnerable to COVID-19 infection. Interestingly, the key regulators of the aforementioned pathways are angiontensin and PPAR family proteins. Hence, angiotensin and PPAR family proteins can be considered as possible therapeutic targets. Contact: [email protected], [email protected] Supplementary information: Supplementary data are available online.
Ashmita Dey, Sagnik Sen 0002, Ujjwal Maulik
Briefings Bioinform.3
2021 Application of active learning in DNA microarray data for cancerous gene identification
Shemim Begum, Ram Sarkar, Debasis Chakraborty, Sagnik Sen 0002, Ujjwal Maulik
Expert Syst. Appl.5
2021 A game theory-based approach to fuzzy clustering for pixel classification in remote sensing imagery
Srimanta Kundu, Ujjwal Maulik, Anirban Mukhopadhyay 0001
Soft Comput.2
2021 Energy-efficient cluster head selection algorithm for IoT using modified glow-worm swarm optimization
Mohana Bakshi, Chandreyee Chowdhury, Ujjwal Maulik
J. Supercomput.3
2020 Special issue on deep learning for video text analysis
Subhadip Basu, Ujjwal Maulik, Umapada Pal 0001
Pattern Recognit. Lett.2
2019 Improved Fuzzy Clustering using Ensemble based Differential Evolution for Remote Sensing Image
abstract
Identification of homogeneous regions in a satellite image is essentially the clustering of pixels in intensity space. Importantly remote sensing image like satellite images contain varieties of land cover types. Some of the covers are significantly large areas whereas some are relatively smaller regions. Therefore, automatically detecting such wide varying areas is a challenging task. Hence, this fact motivated us to propose an improved clustering technique viz. Ensemble based Differential Evolution for Fuzzy Clustering (EDEFC). For this purpose, very recently developed three variants of differential evolution (DE) are used in order to perform the clustering with different set of solutions. As a result, better clustering solution yields from the ensemble of DEs by exhaustive exploration of search space. The proposed EDEFC technique is applied on two numeric remote sensing datasets and Indian Remote Sensing (IRS) satellite image of Kolkata. The results of the EDEFC are shown quantitatively and visually by comparing with eight other clustering techniques. Moreover, the statistical significance test has also been performed in order to judge the superiority of EDEFC.
Jnanendra Prasad Sarkar, Indrajit Saha, Ujjwal Maulik
TENCON3
2019 Identification of infectious disease-associated host genes using machine learning techniques
abstract
BACKGROUND: With the global spread of multidrug resistance in pathogenic microbes, infectious diseases emerge as a key public health concern of the recent time. Identification of host genes associated with infectious diseases will improve our understanding about the mechanisms behind their development and help to identify novel therapeutic targets. RESULTS: We developed a machine learning techniques-based classification approach to identify infectious disease-associated host genes by integrating sequence and protein interaction network features. Among different methods, Deep Neural Networks (DNN) model with 16 selected features for pseudo-amino acid composition (PAAC) and network properties achieved the highest accuracy of 86.33% with sensitivity of 85.61% and specificity of 86.57%. The DNN classifier also attained an accuracy of 83.33% on a blind dataset and a sensitivity of 83.1% on an independent dataset. Furthermore, to predict unknown infectious disease-associated host genes, we applied the proposed DNN model to all reviewed proteins from the database. Seventy-six out of 100 highly-predicted infectious disease-associated genes from our study were also found in experimentally-verified human-pathogen protein-protein interactions (PPIs). Finally, we validated the highly-predicted infectious disease-associated genes by disease and gene ontology enrichment analysis and found that many of them are shared by one or more of the other diseases, such as cancer, metabolic and immune related diseases. CONCLUSIONS: To the best of our knowledge, this is the first computational method to identify infectious disease-associated host genes. The proposed method will help large-scale prediction of host genes associated with infectious-diseases. However, our results indicated that for small datasets, advanced DNN-based method does not offer significant advantage over the simpler supervised machine learning techniques, such as Support Vector Machine (SVM) or Random Forest (RF) for the prediction of infectious disease-associated host genes. Significant overlap of infectious disease with cancer and metabolic disease on disease and gene ontology enrichment analysis suggests that these diseases perturb the functions of the same cellular signaling pathways and may be treated by drugs that tend to reverse these perturbations. Moreover, identification of novel candidate genes associated with infectious diseases would help us to explain disease pathogenesis further and develop novel therapeutics.
Ranjan Kumar Barman, Anirban Mukhopadhyay 0001, Ujjwal Maulik, Santasabuj Das
BMC Bioinform.3
2019 Understanding the evolutionary trend of intrinsically structural disorders in cancer relevant proteins as probed by Shannon entropy scoring and structure network analysis
abstract
BACKGROUND: Malignant diseases have become a threat for health care system. A panoply of biological processes is involved as the cause of these diseases. In order to unveil the mechanistic details of these diseased states, we analyzed protein families relevant to these diseases. RESULTS: Our present study pivots around four apparently unrelated cancer types among which two are commonly occurring viz. Prostate Cancer, Breast Cancer and two relatively less frequent viz. Acute Lymphoblastic Leukemia and Lymphoma. Eight protein families were found to have implications for these cancer types. Our results strikingly reveal that some of the proteins with implications in the cancerous cellular states were showing the structural organization disparate from the signature of the family it constitutes. The sequences were further mapped onto respective structures and compared with the entropic profile. The structures reveal that entropic scores were able to reveal the inherent structural bias of these proteins with quantitative precision, otherwise unseen from other analysis. Subsequently, the betweenness centrality scoring of each residue from the structure network models was resorted to explore the changes in dependencies on residue owing to structural disorder. CONCLUSION: These observations help to obtain the mechanistic changes resulting from the structural orchestration of protein structures. Finally, the hydropathy indexes were obtained to validate the sequence space observations using Shannon entropy and in-turn establishing the compatibility.
Sagnik Sen 0002, Ashmita Dey, Sourav Chowdhury, Ujjwal Maulik, Krishnananda Chattopadhyay
BMC Bioinform.4
2019 Recursive Memetic Algorithm for gene selection in microarray data
Manosij Ghosh, Shemim Begum, Ram Sarkar, Debasis Chakraborty, Ujjwal Maulik
Expert Syst. Appl.5
2019 Integrated Rough Fuzzy Clustering for Categorical data Analysis
Indrajit Saha, Jnanendra Prasad Sarkar, Ujjwal Maulik
Fuzzy Sets Syst.3
2018 Discovering Perturbation of Modular Structure in HIV Progression by Integrating Multiple Data Sources Through Non-Negative Matrix Factorization
abstract
Detecting perturbation in modular structure during HIV-1 disease progression is an important step to understand stage specific infection pattern of HIV-1 virus in human cell. In this article, we proposed a novel methodology on integration of multiple biological information to identify such disruption in human gene module during different stages of HIV-1 infection. We integrate three different biological information: gene expression information, protein-protein interaction information, and gene ontology information in single gene meta-module, through non negative matrix factorization (NMF). As the identified meta-modules inherit those information so, detecting perturbation of these, reflects the changes in expression pattern, in PPI structure and in functional similarity of genes during the infection progression. To integrate modules of different data sources into strong meta-modules, NMF based clustering is utilized here. Perturbation in meta-modular structure is identified by investigating the topological and intramodular properties and putting rank to those meta-modules using a rank aggregation algorithm. We have also analyzed the preservation structure of significant GO terms in which the human proteins of the meta-modules participate. Moreover, we have performed an analysis to show the change of coregulation pattern of identified transcription factors (TFs) over the HIV progression stages.
Sumanta Ray, Ujjwal Maulik
IEEE ACM Trans. Comput. Biol. Bioinform.2
2017 Improving Modified Differential Evolution for Fuzzy Clustering
Jnanendra Prasad Sarkar, Indrajit Saha, Anasua Sarkar, Ujjwal Maulik
HIS4
2017 Multi-size patch based collaborative representation for Palm Dorsa Vein Pattern recognition by enhanced ensemble learning with modified interactive artificial bee colony algorithm
Sandip Joardar, Amitava Chatterjee, Sanghamitra Bandyopadhyay, Ujjwal Maulik
Eng. Appl. Artif. Intell.4
2016 A new evolutionary microRNA marker selection using next-generation sequencing data
abstract
Next-generation sequencing allows high-throughput measurements of non-coding RNA expression levels in tissues. Analysis of microRNAs (miRNAs) is particularly effective in differentiation of cancerous tissue samples, based on patterns of their expression levels. The paper presents a wrapper feature selection approach based on t-Distributed Stochastic Neighbor Embedding (t-SNE), Covariance Matrix Adaptation Evolution Strategy (CMA-Es) and Support Vector Machine (SVM). The advantage of t-SNE is amplification of pairwise similarities by the means of t-Student neighborhood function. The attributes are embedded into 1-D space to reveal similarities between the features. Such information is used by CMA-ES through real-valued encoding in order to model pairwise relations between miRNAs with covariance matrices. Finally, the wrapper uses SVM to evaluate the objective, which expresses the tradeoff between classification quality and the desired number of features. The approach is tested on eight different cancer types from The Cancer Genome Atlas. It allows to find small sets of miRNAs to differentiate cancer types from a single tumor class to the normal one with high certainty.
Adrian Lancucki, Indrajit Saha, Shib Sankar Bhowmick, Ujjwal Maulik, Piotr Lipinski
CEC4
2016 Stability of Consensus Node Orderings Under Imperfect Network Data
abstract
In complex network analysis, the problem of ranking individual nodes based on their importance has attracted increasing attention from the scientific community due to its vast application, such as identification of influential spreaders for viral marketing or epidemic control, bottlenecks for traffic congestion control, and so on. The growing literature proposes a number of measures to determine the rank order of the network entities where complete information about the nodes and their interaction is available. Degree centrality, PageRank, eigenvector centrality, closeness centrality are few such popular measures. In most real-life scenarios, however, the information about the underlying network is incomplete or affected due to noise. The few works that study the effects of incomplete information on the rank orders show the vulnerability of the rank orders in various topologies. In this paper, we investigate the effects of noise, both random and nonrandom, on the aggregated rank orders determined from the degree, PageRank, eigenvector centrality, and closeness centrality-based rankings. This paper reveals an important insight that even the simple Borda Count ranking has the potential to improve on the accuracy of rank orders in networks with uncertainty. This paper shows the existence of stable nodes in various networks and indicates that the design of the consensus approach based on the properties of the stable nodes can further improve the stability of the rank orders.
Srinka Basu, Ujjwal Maulik, Oishik Chatterjee
IEEE Trans. Comput. Soc. Syst.2
2016 A Game Theory Inspired Approach to Stable Core Decomposition on Weighted Networks
abstract
Meso-scale structural analysis, like core decomposition has uncovered groups of nodes that play important roles in the underlying complex systems. The existing core decomposition approaches generally focus on node properties like degree and strength. The node centric approaches can only capture a limited information about the local neighborhood topology. In the present work, we propose a group density based core analysis approach that overcome the drawbacks of the node centric approaches. The proposed algorithmic approach focuses on weight density, cohesiveness, and stability of a substructure. The method also assigns an unique score to every node that rank the nodes based on their degree of core-ness. To determine the correctness of the proposed method, we propose a synthetic benchmark with planted core structure. A performance test on the null model is carried out using a weighted lattice without core structures. We further test the stability of the approach against random noise. The experimental results prove the superiority of our algorithm over the state-of-the-arts. We finally analyze the core structures of several popular weighted network models and real life weighted networks. The experimental results reveal important node ranking and hierarchical organization of the complex networks, which give us better insight about the underlying systems.
Srinka Basu, Ujjwal Maulik
IEEE Trans. Knowl. Data Eng.2
2015 A review of in silico approaches for analysis and prediction of HIV-1-human protein-protein interactions
abstract
The computational or in silico approaches for analysing the HIV-1-human protein-protein interaction (PPI) network, predicting different host cellular factors and PPIs and discovering several pathways are gaining popularity in the field of HIV research. Although there exist quite a few studies in this regard, no previous effort has been made to review these works in a comprehensive manner. Here we review the computational approaches that are devoted to the analysis and prediction of HIV-1-human PPIs. We have broadly categorized these studies into two fields: computational analysis of HIV-1-human PPI network and prediction of novel PPIs. We have also presented a comparative assessment of these studies and proposed some methodologies for discussing the implication of their results. We have also reviewed different computational techniques for predicting HIV-1-human PPIs and provided a comparative study of their applicability. We believe that our effort will provide helpful insights to the HIV research community.
Sanghamitra Bandyopadhyay, Sumanta Ray, Anirban Mukhopadhyay 0001, Ujjwal Maulik
Briefings Bioinform.4
2015 Priority based ∈ dominance: A new measure in multiobjective optimization
Sanghamitra Bandyopadhyay, Rudrasis Chakraborty, Ujjwal Maulik
Inf. Sci.3
2015 MiRNA-TF-gene network analysis through ranking of biomolecules for multi-informative uterine leiomyoma dataset
Saurav Mallik, Ujjwal Maulik
J. Biomed. Informatics2
2015 Ensemble based rough fuzzy clustering for categorical data
Indrajit Saha, Jnanendra Prasad Sarkar, Ujjwal Maulik
Knowl. Based Syst.3
2015 Predicting Protein-Protein Interaction Sites with a Novel Membership Based Fuzzy SVM Classifier
abstract
Predicting residues that participate in protein-protein interactions (PPI) helps to identify, which amino acids are located at the interface. In this paper, we show that the performance of the classical support vector machine (SVM) algorithm can further be improved with the use of a custom-designed fuzzy membership function, for the partner-specific PPI interface prediction problem. We evaluated the performances of both classical SVM and fuzzy SVM (F-SVM) on the PPI databases of three different model proteomes of Homo sapiens, Escherichia coli and Saccharomyces Cerevisiae and calculated the statistical significance of the developed F-SVM over classical SVM algorithm. We also compared our performance with the available state-of-the-art fuzzy methods in this domain and observed significant performance improvements. To predict interaction sites in protein complexes, local composition of amino acids together with their physico-chemical characteristics are used, where the F-SVM based prediction method exploits the membership function for each pair of sequence fragments. The average F-SVM performance (area under ROC curve) on the test samples in 10-fold cross validation experiment are measured as 77.07, 78.39, and 74.91 percent for the aforementioned organisms respectively. Performances on independent test sets are obtained as 72.09, 73.24 and 82.74 percent respectively. The software is available for free download from http://code.google.com/p/cmater-bioinfo.
Brijesh Kumar Sriwastava, Subhadip Basu, Ujjwal Maulik
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 Nonintrusive Load Monitoring: A Temporal Multilabel Classification Approach
abstract
The article tackles the issues related to the identification of electrical appliances inside residential buildings. Each appliance can be identified from the aggregate power readings at the meter panel. The possibility of applying a temporal multilabel classification approach in the domain of nonintrusive load monitoring is explored (nonevent-based method). A novel set of metafeatures is proposed. The method is tested on sampling rates based on the capabilities of current smart meters. The proposed approach is validated over a dataset of energy readings at residences for a period of a year for 100 houses containing different sets of appliances (water heater, washing machines, etc.). This method is applicable for the demand side management of households in the current limitation of smart meters; from the inhabitants or from the grid operator's point of view.
Kaustav Basu, Vincent Debusschere, Seddik Bacha, Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE Trans. Ind. Informatics4
2014 Incorporating the type and direction information in predicting novel regulatory interactions between HIV-1 and human proteins using a biclustering approach
abstract
BACKGROUND: Discovering novel interactions between HIV-1 and human proteins would greatly contribute to different areas of HIV research. Identification of such interactions leads to a greater insight into drug target prediction. Some recent studies have been conducted for computational prediction of new interactions based on the experimentally validated information stored in a HIV-1-human protein-protein interaction database. However, these techniques do not predict any regulatory mechanism between HIV-1 and human proteins by considering interaction types and direction of regulation of interactions. RESULTS: Here we present an association rule mining technique based on biclustering for discovering a set of rules among human and HIV-1 proteins using the publicly available HIV-1-human PPI database. These rules are subsequently utilized to predict some novel interactions among HIV-1 and human proteins. For prediction purpose both the interaction types and direction of regulation of interactions, (i.e., virus-to-host or host-to-virus) are considered here to provide important additional information about the regulation pattern of interactions. We have also studied the biclusters and analyzed the significant GO terms and KEGG pathways in which the human proteins of the biclusters participate. Moreover the predicted rules have also been analyzed to discover regulatory relationship between some human proteins in course of HIV-1 infection. Some experimental evidences of our predicted interactions have been found by searching the recent literatures in PUBMED. We have also highlighted some human proteins that are likely to act against the HIV-1 attack. CONCLUSIONS: We pose the problem of identifying new regulatory interactions between HIV-1 and human proteins based on the existing PPI database as an association rule mining problem based on biclustering algorithm. We discover some novel regulatory interactions between HIV-1 and human proteins. Significant number of predicted interactions has been found to be supported by recent literature.
Anirban Mukhopadhyay 0001, Sumanta Ray, Ujjwal Maulik
BMC Bioinform.3
2014 Incremental learning based multiobjective fuzzy clustering for categorical data
Indrajit Saha, Ujjwal Maulik
Inf. Sci.2
2014 Multi-level thresholding using quantum inspired meta-heuristics
Sandip Dey, Indrajit Saha, Siddhartha Bhattacharyya 0001, Ujjwal Maulik
Knowl. Based Syst.4
2014 Integration of dense subgraph finding with feature clustering for unsupervised feature selection
Sanghamitra Bandyopadhyay, Tapas Bhadra, Pabitra Mitra, Ujjwal Maulik
Pattern Recognit. Lett.4
2014 A Survey of Multiobjective Evolutionary Algorithms for Data Mining: Part I
abstract
The aim of any data mining technique is to build an efficient predictive or descriptive model of a large amount of data. Applications of evolutionary algorithms have been found to be particularly useful for automatic processing of large quantities of raw noisy data for optimal parameter setting and to discover significant and meaningful information. Many real-life data mining problems involve multiple conflicting measures of performance, or objectives, which need to be optimized simultaneously. Under this context, multiobjective evolutionary algorithms are gradually finding more and more applications in the domain of data mining since the beginning of the last decade. In this two-part paper, we have made a comprehensive survey on the recent developments of multiobjective evolutionary algorithms for data mining problems. In this paper, Part I, some basic concepts related to multiobjective optimization and data mining are provided. Subsequently, various multiobjective evolutionary approaches for two major data mining tasks, namely feature selection and classification, are surveyed. In Part II of this paper, we have surveyed different multiobjective evolutionary algorithms for clustering, association rule mining, and several other data mining tasks, and provided a general discussion on the scopes for future research in this domain.
Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Carlos A. Coello Coello
IEEE Trans. Evol. Comput.2
2014 Survey of Multiobjective Evolutionary Algorithms for Data Mining: Part II
abstract
This paper is the second part of a two-part paper, which is a survey of multiobjective evolutionary algorithms for data mining problems. In Part I , multiobjective evolutionary algorithms used for feature selection and classification have been reviewed. In this part, different multiobjective evolutionary algorithms used for clustering, association rule mining, and other data mining tasks are surveyed. Moreover, a general discussion is provided along with scopes for future research in the domain of multiobjective evolutionary algorithms for data mining.
Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Carlos A. Coello Coello
IEEE Trans. Evol. Comput.2
2014 Guest Editorial: Special Issue on Advances in Multiobjective Evolutionary Algorithms for Data Mining
abstract
The six articles in this special issue provide a snapshot of the current research trends in multiobjective evolutionary algorithms for data mining. The main issues and challenges in this domain have been highlighted and directions of future research work have also been provided,
Sanghamitra Bandyopadhyay, Ujjwal Maulik, Carlos A. Coello Coello, Witold Pedrycz
IEEE Trans. Evol. Comput.2
2013 Integrated analysis of gene expression and genome-wide DNA methylation for tumor prediction: An association rule mining-based approach
abstract
Statistical analysis and association rule mining are two most efficient techniques, where the first one is used to identify differentially expressed/methylated genes across different types of samples or experimental conditions and the second one is used to determine expression/methylation relationships among them. In this article, we have performed an integrated analysis of statistical methods and association rule mining on mRNA expression and DNA methylation datasets for the prediction of Uterine Leiomyoma. Moreover, we have proposed a novel rule-base classifier. Depending on 16 different rule-interestingness measures, we have applied a Genetic Algorithm based rank aggregation technique on the association rules which are generated from the training data by Apriori association rule mining algorithm. After determining the ranks of the rules, we have conducted a majority voting technique on each test point to determine its class-label (i.e. tumor or normal class-label) through weighted-sum method. We have run this classifier on the combined dataset using k-fold cross-validation and also performed a comparative performance analysis with other popular rule-base classifiers. Finally, we have predicted the status of some important genes (through frequency analysis in association rules for tumor and normal class-labels individually) that have a major role for tumor formation in Uterine Leiomyoma.
Saurav Mallik, Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay
CIBCB3
2013 Fuzzy rule-based classifier for microarray gene expression data by using a multiobjective PSO-based approach
abstract
In this article a fuzzy rule-based classifier has been designed on the framework of multiobjective Particle Swarm Optimization. The proposed approach is applied on microarray gene expression data to obtain genes with significant expression with respect to two different classes. Two fuzzy sets are represented with the linguistic values “high” and “low”. On the training dataset, the proposed approach is applied for the purpose of deriving good classification rules. To be precise, the good rules are those that have less attributes in the antecedent part and provide maximum accuracy. Moreover we also consider the existing redundancy among the selected rules which should be minimized. Here the underlying structure is modeled using multiobjective PSO with the support of non-dominated sorting and crowding distance sorting. The first objective is to maximize the classification accuracy and second objective is to minimize the rule-base complexity (number of rules and average rule length) and the redundancy of the rules. The performance of the proposed algorithm is compared with that of single objective versions, Support Vector machine classifier and Bayes classifier on several real-life datasets.
Monalisa Mandal, Anirban Mukhopadhyay 0001, Ujjwal Maulik
FUZZ-IEEE3
2013 Incorporating fuzzy semantic similarity measure in detecting human protein complexes in PPI network: A multiobjective approach
abstract
Detection of protein complexes within protein-protein interaction networks (PPIN) is a valuable step toward the analysis of biological processes and pathways. Several high-throughput experimental techniques produce large number of PPIs that can be extensively utilized for constructing PPI network of a species. Decomposition of the whole PPI network into smaller and manageable modules is an ongoing challenge. Here we have developed a multi-objective algorithm for detecting human protein complexes by partitioning large human PPI network into clusters which serve as protein complexes. Some graphical properties like density, centrality etc., are utilized for building the objectives. Besides the graphical properties we have also exploited a fuzzy measure based semantic similarity approach to construct similarity based objective. The proposed technique is demonstrated in the human PPI network and the resulting complexes are analyzed in context of Gene Ontology (GO) and pathway enrichment. We have also compared our results with that of some state-of-the-art algorithms in context of different performance metrics. The biological relevance of our predicted complexes are also established here by linking them with 22 key disease classes.
Sumanta Ray, Sanghamitra Bandyopadhyay, Anirban Mukhopadhyay 0001, Ujjwal Maulik
FUZZ-IEEE4
2013 Mining Quasi-Bicliques from HIV-1-Human Protein Interaction Network: A Multiobjective Biclustering Approach
abstract
In this work, we model the problem of mining quasi-bicliques from weighted viral-host protein-protein interaction network as a biclustering problem for identifying strong interaction modules. In this regard, a multiobjective genetic algorithm-based biclustering technique is proposed that simultaneously optimizes three objective functions to obtain dense biclusters having high mean interaction strengths. The performance of the proposed technique has been compared with that of other existing biclustering methods on an artificial data. Subsequently, the proposed biclustering method is applied on the records of biologically validated and predicted interactions between a set of HIV-1 proteins and a set of human proteins to identify strong interaction modules. For this, the entire interaction information is realized as a bipartite graph. We have further investigated the biological significance of the obtained biclusters. The human proteins involved in the strong interaction module have been found to share common biological properties and they are identified as the gateways of viral infection leading to various diseases. These human proteins can be potential drug targets for developing anti-HIV drugs.
Ujjwal Maulik, Anirban Mukhopadhyay 0001, Malay Bhattacharyya 0001, Lars Kaderali, Benedikt Brors, Sanghamitra Bandyopadhyay, Roland Eils
IEEE ACM Trans. Comput. Biol. Bioinform.1
2013 Reformulated Kemeny Optimal Aggregation with Application in Consensus Ranking of microRNA Targets
abstract
MicroRNAs are very recently discovered small noncoding RNAs, responsible for negative regulation of gene expression. Members of this endogenous family of small RNA molecules have been found implicated in many genetic disorders. Each microRNA targets tens to hundreds of genes. Experimental validation of target genes is a time- and cost-intensive procedure. Therefore, prediction of microRNA targets is a very important problem in computational biology. Though, dozens of target prediction algorithms have been reported in the past decade, they disagree significantly in terms of target gene ranking (based on predicted scores). Rank aggregation is often used to combine multiple target orderings suggested by different algorithms. This technique has been used in diverse fields including social choice theory, meta search in web, and most recently, in bioinformatics. Kemeny optimal aggregation (KOA) is considered the more profound objective for rank aggregation. The consensus ordering obtained through Kemeny optimal aggregation incurs minimum pairwise disagreement with the input orderings. Because of its computational intractability, heuristics are often formulated to obtain a near optimal consensus ranking. Unlike its real time use in meta search, there are a number of scenarios in bioinformatics (e.g., combining microRNA target rankings, combining disease-related gene rankings obtained from microarray experiments) where evolutionary approaches can be afforded with the ambition of better optimization. We conjecture that an ideal consensus ordering should have its total disagreement shared, as equally as possible, with the input orderings. This is also important to refrain the evolutionary processes from getting stuck to local extremes. In the current work, we reformulate Kemeny optimal aggregation while introducing a trade-off between the total pairwise disagreement and its distribution. A simulated annealing-based implementation of the proposed objective has been found effective in context of microRNA target ranking. Supplementary data and source code link are available at: >http://www.isical.ac.in/bioinfo_miu/ieee_tcbb_kemeny.rar.
Debarka Sengupta, Aroonalok Pyne, Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE ACM Trans. Comput. Biol. Bioinform.3
2012 Score Based Aggregation of microRNA Target Orderings
Debarka Sengupta, Ujjwal Maulik, Sanghamitra Bandyopadhyay
ISBRA2
2012 δ-TRIMAX: Extracting Triclusters and Analysing Coregulation in Time Series Gene Expression Data
Anirban Bhar, Martin Haubrock, Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Edgar Wingender
WABI4
2012 Efficient parallel algorithm for pixel classification in remote sensing imagery
Ujjwal Maulik, Anasua Sarkar
GeoInformatica1
2012 A parallel bi-directional self-organizing neural network (PBDSONN) architecture for color image extraction and segmentation
Siddhartha Bhattacharyya 0001, Ujjwal Maulik, Paramartha Dutta
Neurocomputing2
2012 SVMeFC: SVM Ensemble Fuzzy Clustering for Satellite Image Segmentation
abstract
The problem of unsupervised image segmentation of a satellite image in a number of homogeneous regions can be viewed as the task of clustering the pixels in the intensity space. This letter presents an approach that exploits the capability of some recently proposed fuzzy clustering techniques, as well as support vector machine (SVM) classifiers, to yield improved solutions. All the fuzzy clustering techniques are first used to produce a set of different clustering solutions. Each such solution has been improved by a novel technique based on an SVM classifier. Thereafter, the cluster-based similarity partition algorithm is used to create the final clustering solution from all improved ensemble solutions. Results demonstrating the effectiveness of the proposed technique are provided for numeric remote sensing data described in terms of feature vectors. Moreover, a remotely sensed image of Calcutta City has been segmented using the proposed technique to establish its utility. In addition, the additional information of this letter is given as supplementary at http://sysbio.icm.edu.pl/indra/SVMeFC.html.
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
IEEE Geosci. Remote. Sens. Lett.2
2012 Weighted Markov Chain Based Aggregation of Biomolecule Orderings
abstract
The scope and effectiveness of Rank Aggregation (RA) have already been established in contemporary bioinformatics research. Rank aggregation helps in meta-analysis of putative results collected from different analytic or experimental sources. For example, we often receive considerably differing ranked lists of genes or microRNAs from various target prediction algorithms or microarray studies. Sometimes combining them all, in some sense, yields more effective ordering of the set of objects. Also, assigning a certain level of confidence to each source of ranking is a natural demand of aggregation. Assignment of weights to the sources of orderings can be performed by experts. Several rank aggregation approaches like those based on Markov Chains (MCs), evolutionary algorithms, etc., exist in the literature. Markov chains, in general, are faster than the evolutionary approaches. Unlike the evolutionary computing approaches Markov chains have not been used for weighted aggregation scenarios. This is because of the absence of a formal framework of Weighted Markov Chain (WMC). In this paper, we propose the use of a modified version of MC4 (one of the Markov chains proposed by Dwork et al., 2001), followed by the weighted analog of local Kemenization for performing rank aggregation, where the sources of rankings can be prioritized by an expert. Effectiveness of the weighted Markov chain approach over the very recently proposed Genetic Algorithm (GA) and Cross-Entropy Monte Carlo (MC) algorithm-based techniques, has been established for gene orderings from microarray analysis and orderings of predicted microRNA targets.
Debarka Sengupta, Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE ACM Trans. Comput. Biol. Bioinform.2
2011 Discovery of MicroRNA markers: An SVM-based multiobjective feature selection approach
abstract
MicroRNAs (miRNAs) are small non-coding RNAs that have been shown to play important roles in gene regulation and various biological processes. The abnormal expression of some specific miRNAs often results in the development of cancer. In this article, we have utilized a multiobjective genetic algorithm-based feature selection algorithm wrapped with support vector machine (SVM) classifier for selecting promising miRNAs having differential expression in benign and malignant tissue samples. Subsequently, the non-dominated sets of promising miRNAs are aggregated into a single most promising miRNA subset. Finally, the Signal-to-Noise Ratio (SNR) statistic has been applied on the obtained miRNA subset for identifying potential miRNA markers that distinguish the two classes (benign and malignant) of tissue samples. The performance has been demonstrated on four real-life miRNA expression datasets for different SVM kernel functions and the identified miRNA markers are reported.
Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay
CIBCB2
2011 PMAFC: A New Probabilistic Memetic Algorithm Based Fuzzy Clustering
Indrajit Saha, Ujjwal Maulik, Dariusz Plewczynski
ISMIS2
2011 Improvement of new automatic differential fuzzy clustering using SVM classifier for microarray analysis
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
Expert Syst. Appl.2
2011 Unsupervised and Supervised Learning Approaches Together for Microarray Analysis
abstract
In this article, a novel concept is introduced by using both unsupervised and supervised learning. For unsupervised learning, the problem of fuzzy clustering in microarray data as a multiobjective optimization is used, which simultaneously optimizes two internal fuzzy cluster validity indices to yield a set of Pareto-optimal clustering solutions. In this regards, a new multiobjective differential evolution based fuzzy clustering technique has been proposed. Subsequently, for supervised learning, a fuzzy majority voting scheme along with support vector machine is used to integrate the clustering information from all the solutions in the resultant Pareto-optimal set. The performances of the proposed clustering techniques have been demonstrated on five publicly available benchmark microarray data sets. A detail comparison has been carried out with multiobjective genetic algorithm based fuzzy clustering, multiobjective differential evolution based fuzzy clustering, single objective versions of differential evolution and genetic algorithm based fuzzy clustering as well as well known fuzzy c-means algorithm. While using support vector machine, comparative studies of the use of four different kernel functions are also reported. Statistical significance test has been done to establish the statistical superiority of the proposed multiobjective clustering approach. Finally, biological significance test has been carried out using a web based gene annotation tool to show that the proposed integrated technique is able to produce biologically relevant clusters of coexpressed genes.
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
Fundam. Informaticae2
2011 A self-trained ensemble with semisupervised SVM: An application to pixel classification of remote sensing imagery
Ujjwal Maulik, Debasis Chakraborty
Pattern Recognit.1
2010 Simultaneous informative gene selection and clustering through multiobjective optimization
abstract
Clustering methods are used for unsupervised classification of tumor subclasses in microarray gene expression data sets organized in a fashion where the rows represent the tumor samples and columns represent the genes. Clustering algorithms can be very sensitive with respect to the set of features (genes) considered in the clustering process. It is important to select the set of informative and relevant genes to be used for clustering. In this article, a multiobjective genetic algorithm based technique has been proposed for performing the tasks of gene selection and fuzzy clustering simultaneously. A novel encoding technique is developed in this regard and the algorithm searches for the best cluster centers while minimizing the number of selected genes. The number of clusters is evolved automatically. The performance of the proposed technique has been illustrated on an artificial data set and compared with that of several other related feature selection/clustering approaches. Moreover its performance is demonstrated on two real life multi-class gene expression data sets viz., Brain tumor and Lung tumor data sets.
Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE Congress on Evolutionary Computation2
2010 Real-coded differential crisp clustering for MRI brain image segmentation
abstract
In this paper, a segmentation technique of multi-spectral magnetic resonance image of the brain using a new differential evolution based crisp clustering is proposed. Real-coded encoding of the cluster centres is used for this purpose. Here assignments of points to different clusters are made based on the Euclidean distance. The proposed method is applied on several simulated T1-weighted, T2-weighted and proton density for normal and MS lesion magnetic resonance brain images. Superiority of the proposed method over genetic algorithm based crisp clustering, simulated annealing based crisp clustering, K-means and average linkage are demonstrated quantitatively. Segmentation obtained by differential evolution based crisp clustering technique is also compared with the available ground truth information. Also statistical analysis has been conducted to judge the effectiveness. Matlab version of the software is available at http://bio.icm.edu.pl/~darman/MRI.
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
IEEE Congress on Evolutionary Computation2
2010 SFSSClass: an integrated approach for miRNA based tumor classification
abstract
BACKGROUND: MicroRNA (miRNA) expression profiling data has recently been found to be particularly important in cancer research and can be used as a diagnostic and prognostic tool. Current approaches of tumor classification using miRNA expression data do not integrate the experimental knowledge available in the literature. A judicious integration of such knowledge with effective miRNA and sample selection through a biclustering approach could be an important step in improving the accuracy of tumor classification. RESULTS: In this article, a novel classification technique called SFSSClass is developed that judiciously integrates a biclustering technique SAMBA for simultaneous feature (miRNA) and sample (tissue) selection (SFSS), a cancer-miRNA network that we have developed by mining the literature of experimentally verified cancer-miRNA relationships and a classifier uncorrelated shrunken centroid (USC). SFSSClass is used for classifying multiple classes of tumors and cancer cell lines. In a part of the investigation, poorly differentiated tumors (PDT) having non diagnostic histological appearance are classified while training on more differentiated tumor (MDT) samples. The proposed method is found to outperform the best known accuracy in the literature on the experimental data sets. For example, while the best accuracy reported in the literature for classifying PDT samples is approximately 76.5%, the accuracy of SFSSClass is found to be approximately 82.3%. The advantage of incorporating biclustering integrated with the cancer-miRNA network is evident from the consistently better performance of SFSSClass (integration of SAMBA, cancer-miRNA network and USC) over USC (eg., approximately 70.5% for SFSSClass versus approximately 58.8% in classifying a set of 17 MDT samples from 9 tumor types, approximately 91.7% for SFSSClass versus approximately 75% in classifying 12 cell lines from 6 tumor types and approximately 82.3% for SFSSClass versus approximately 41.2% in classifying 17 PDT samples from 11 tumor types). CONCLUSION: In this article, we develop the SFSSClass algorithm which judiciously integrates a biclustering technique for simultaneous feature (miRNA) and sample (tissue) selection, the cancer-miRNA network and a classifier. The novel integration of experimental knowledge with computational tools efficiently selects relevant features that have high intra-class and low inter-class similarity. The performance of the SFSSClass is found to be significantly improved with respect to the other existing approaches.
Ramkrishna Mitra, Sanghamitra Bandyopadhyay, Ujjwal Maulik, Michael Q. Zhang
BMC Bioinform.3
2010 A Robust Multiple Classifier System for Pixel Classification of Remote Sensing Images
abstract
Satellite image classification is a complex process that may be affected by many factors. This article addresses the problem of pixel classification of satellite images by a robust multiple classifier system that combines k-NN, support vector machine (SVM) and incremental learning algorithm (IL). The effectiveness of this combination is investigated for satellite imagery which usually have overlapping class boundaries. These classifiers are initially designed using a small set of labeled points. Combination of these algorithms has been done based on majority voting rule. The effectiveness of the proposed technique is first demonstrated for a numeric remote sensing data described in terms of feature vectors and then identifying different land cover regions in remote sensing imagery. Experimental results on numeric data as well as two remote sensing data show that employing combination of classifiers can effectively increase the accuracy label. Comparison is made with each of these single classifiers in terms of kappa value, accuracy, cluster quality indices and visual quality of the classified images.
Ujjwal Maulik, Debasis Chakraborty
Fundam. Informaticae1
2010 Evolutionary Rough Parallel Multi-Objective Optimization Algorithm
abstract
A hybrid unsupervised learning algorithm, which is termed as Parallel Rough-based Archived Multi-Objective Simulated Annealing (PARAMOSA), is proposed in this article. It comprises a judicious integration of the principles of the rough sets theory and the scalable distributed paradigm with the archived multi-objective simulated annealing approach. While the concept of boundary approximations of rough sets in this implementation, deals with the incompleteness in the dynamic classification method with the quality of classification coefficient as the classificatory competencemeasurement, the time-efficient parallel approach enables faster convergence of the Pareto-archived evolution strategy. It incorporates both the rough set-based dynamic archive classifi- cation method and the distributed implementation as a two-phase speedup strategy in this algorithm. A measure of the amount of domination between two solutions has been incorporated in this work to determine the acceptance probability of a new solution with an improvement in the spread of the non-dominated solutions in the Pareto-front by adopting rough sets theory. A complexity analysis of the proposed algorithm is provided. An extensive comparative study of the proposed algorithm with three other existing and well-known Multi-Objective Evolutionary Algorithms (MOEAs) demonstrate the effectiveness of the former with respect to four existing performance metrics and eleven benchmark test problems of varying degrees of difficulties. The superiority of this new parallel implementation over other algorithms also has been demonstrated in timing, which achieves a near optimal speedup with a minimal communication overhead.
Ujjwal Maulik, Anasua Sarkar
Fundam. Informaticae1
2010 Automatic Fuzzy Clustering Using Modified Differential Evolution for Image Classification
abstract
The problem of classifying an image into different homogeneous regions is viewed as the task of clustering the pixels in the intensity space. In particular, satellite images contain landcover types, some of which cover significantly large areas while some (e.g., bridges and roads) occupy relatively much smaller regions. Automatically detecting regions or clusters of such widely varying sizes is a challenging task. In this paper, a new real-coded modified differential evolution based automatic fuzzy clustering algorithm is proposed which automatically evolves the number of clusters as well as the proper partitioning from a data set. Here, the assignment of points to different clusters is done based on a Xie-Beni index where the Euclidean distance is taken into consideration. The effectiveness of the proposed technique is first demonstrated for two numeric remote sensing data described in terms of feature vectors and then in identifying different landcover regions in remote sensing imagery. The superiority of the new method is demonstrated by comparing it with other existing techniques like automatic clustering using improved differential evolution, classical differential evolution based automatic fuzzy clustering, variable length genetic algorithm based fuzzy clustering, and well known fuzzy C-means algorithm both qualitatively and quantitatively.
Ujjwal Maulik, Indrajit Saha
IEEE Trans. Geosci. Remote. Sens.1
2010 Integrating Clustering and Supervised Learning for Categorical Data Analysis
abstract
The problem of fuzzy clustering of categorical data, where no natural ordering among the elements of a categorical attribute domain can be found, is an important problem in exploratory data analysis. As a result, a few clustering algorithms with focus on categorical data have been proposed. In this paper, a modified differential evolution (DE)-based fuzzy c-medoids (FCMdd) clustering of categorical data has been proposed. The algorithm combines both local as well as global information with adaptive weighting. The performance of the proposed method has been compared with those using genetic algorithm, simulated annealing, and the classical DE technique, besides the FCMdd, fuzzy k-modes, and average linkage hierarchical clustering algorithm for four artificial and four real life categorical data sets. Statistical test has been carried out to establish the statistical significance of the proposed method. To improve the result further, the clustering method is integrated with a support vector machine (SVM), a well-known technique for supervised learning. A fraction of the data points selected from different clusters based on their proximity to the respective medoids is used for training the SVM. The clustering assignments of the remaining points are thereafter determined using the trained classifier. The superiority of the integrated clustering and supervised learning approach has been demonstrated.
Ujjwal Maulik, Sanghamitra Bandyopadhyay, Indrajit Saha
IEEE Trans. Syst. Man Cybern. Part A1
2009 Analysis of microarray data using multiobjective variable string length genetic fuzzy clustering
abstract
In this article, a novel multiobjective variable string length real coded genetic fuzzy clustering scheme for clustering microarray gene expression data has been proposed. The proposed technique automatically evolves the number of clusters along with the clustering result. The multiobjective variable string length clustering technique encodes the cluster centers in its chromosomes and simultaneously optimizes two fuzzy validity indices namely PBM index and Xie-Beni validity measure. In the final generation, it produces a set of non-dominated solutions, from which the best solution is selected using Silhouette index which is independent of the number of clusters. The corresponding chromosome length provides the number of clusters. The proposed method is applied on three publicly available real life gene expression data. Superiority of the proposed method over some other well known clustering algorithms has been demonstrated quantitatively.
Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Ujjwal Maulik
IEEE Congress on Evolutionary Computation3
2009 Unsupervised cancer classification through SVM-boosted multiobjective fuzzy clustering with majority voting ensemble
abstract
In this article, we have presented an unsupervised cancer classification technique based on multiobjective genetic fuzzy clustering of the tissue samples. In this regard, coordinate of the cluster centers have been encoded in the chromosomes and three fuzzy cluster validity indices are simultaneously optimized. Each solution of the resultant Pareto-optimal set has been boosted by a novel technique based on Support Vector Machine (SVM) classification. Finally, the clustering information possessed by the non-dominated solutions are combined through a majority voting ensemble technique to produce the final clustering solution. The performance of the proposed multiobjective clustering method has been compared to several other microarray clustering algorithms for three publicly available benchmark cancer data sets, viz., Leukemia, Colon cancer and Lymphoma data to establish its superiority.
Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE Congress on Evolutionary Computation2
2009 Combining Pareto-optimal clusters using supervised learning for identifying co-expressed genes
abstract
BACKGROUND: The landscape of biological and biomedical research is being changed rapidly with the invention of microarrays which enables simultaneous view on the transcription levels of a huge number of genes across different experimental conditions or time points. Using microarray data sets, clustering algorithms have been actively utilized in order to identify groups of co-expressed genes. This article poses the problem of fuzzy clustering in microarray data as a multiobjective optimization problem which simultaneously optimizes two internal fuzzy cluster validity indices to yield a set of Pareto-optimal clustering solutions. Each of these clustering solutions possesses some amount of information regarding the clustering structure of the input data. Motivated by this fact, a novel fuzzy majority voting approach is proposed to combine the clustering information from all the solutions in the resultant Pareto-optimal set. This approach first identifies the genes which are assigned to some particular cluster with high membership degree by most of the Pareto-optimal solutions. Using this set of genes as the training set, the remaining genes are classified by a supervised learning algorithm. In this work, we have used a Support Vector Machine (SVM) classifier for this purpose. RESULTS: The performance of the proposed clustering technique has been demonstrated on five publicly available benchmark microarray data sets, viz., Yeast Sporulation, Yeast Cell Cycle, Arabidopsis Thaliana, Human Fibroblasts Serum and Rat Central Nervous System. Comparative studies of the use of different SVM kernels and several widely used microarray clustering techniques are reported. Moreover, statistical significance tests have been carried out to establish the statistical superiority of the proposed clustering approach. Finally, biological significance tests have been carried out using a web based gene annotation tool to show that the proposed method is able to produce biologically relevant clusters of co-expressed genes. CONCLUSION: The proposed clustering method has been shown to perform better than other well-known clustering algorithms in finding clusters of co-expressed genes efficiently. The clusters of genes produced by the proposed technique are also found to be biologically significant, i.e., consist of genes which belong to the same functional groups. This indicates that the proposed clustering method can be used efficiently to identify co-expressed genes in microarray gene expression data.Supplementary Website The pre-processed and normalized data sets, the matlab code and other related materials are available at http://anirbanmukhopadhyay.50webs.com/mogasvm.html.
Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay
BMC Bioinform.1
2009 Modified differential evolution based fuzzy clustering for pixel classification in remote sensing imagery
Ujjwal Maulik, Indrajit Saha
Pattern Recognit.1
2009 Towards improving fuzzy clustering using support vector machine: Application to gene expression data
Anirban Mukhopadhyay 0001, Ujjwal Maulik
Pattern Recognit.2
2009 Multiobjective Genetic Algorithm-Based Fuzzy Clustering of Categorical Attributes
abstract
Recently, the problem of clustering categorical data, where no natural ordering among the elements of a categorical attribute domain can be found, has been gaining significant attention from researchers. With the growing demand for categorical data clustering, a few clustering algorithms with focus on categorical data have recently been developed. However, most of these methods attempt to optimize a single measure of the clustering goodness. Often, such a single measure may not be appropriate for different kinds of datasets. Thus, consideration of multiple, often conflicting, objectives appears to be natural for this problem. Although we have previously addressed the problem of multiobjective fuzzy clustering for continuous data, these algorithms cannot be applied for categorical data where the cluster means are not defined. Motivated by this, in this paper a multiobjective genetic algorithm-based approach for fuzzy clustering of categorical data is proposed that encodes the cluster modes and simultaneously optimizes fuzzy compactness and fuzzy separation of the clusters. Moreover, a novel method for obtaining the final clustering solution from the set of resultant Pareto-optimal solutions in proposed. This is based on majority voting among Pareto front solutions followed byk-nn classification. The performance of the proposed fuzzy categorical data-clustering techniques has been compared with that of some other widely used algorithms, both quantitatively and qualitatively. For this purpose, various synthetic and real-life categorical datasets have been considered. Also, a statistical significance test has been conducted to establish the significant superiority of the proposed multiobjective approach.
Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE Trans. Evol. Comput.2
2009 Unsupervised Pixel Classification in Satellite Imagery Using Multiobjective Fuzzy Clustering Combined With SVM Classifier
abstract
The problem of unsupervised classification of a satellite image in a number of homogeneous regions can be viewed as the task of clustering the pixels in the intensity space. This paper proposes a novel approach that combines a recently proposed multiobjective fuzzy clustering scheme with support vector machine (SVM) classifier to yield improved solutions. The multiobjective technique is first used to produce a set of nondominated solutions. The nondominated set is then used to find some high-confidence points using a fuzzy voting technique. The SVM classifier is thereafter trained by these high-confidence points. Finally, the remaining points are classified using the trained classifier. Results demonstrating the effectiveness of the proposed technique are provided for numeric remote sensing data described in terms of feature vectors. Moreover, two remotely sensed images of Bombay and Calcutta cities have been classified using the proposed technique to establish its utility.
Anirban Mukhopadhyay 0001, Ujjwal Maulik
IEEE Trans. Geosci. Remote. Sens.2
2009 Medical Image Segmentation Using Genetic Algorithms
abstract
Genetic algorithms (GAs) have been found to be effective in the domain of medical image segmentation, since the problem can often be mapped to one of search in a complex and multimodal landscape. The challenges in medical image segmentation arise due to poor image contrast and artifacts that result in missing or diffuse organ/tissue boundaries. The resulting search space is therefore often noisy with a multitude of local optima. Not only does the genetic algorithmic framework prove to be effective in coming out of local optima, it also brings considerable flexibility into the segmentation procedure. In this paper, an attempt has been made to review the major applications of GAs to the domain of medical image segmentation.
Ujjwal Maulik
IEEE Trans. Inf. Technol. Biomed.1
2009 Finding Multiple Coherent Biclusters in Microarray Data Using Variable String Length Multiobjective Genetic Algorithm
abstract
Microarray technology enables the simultaneous monitoring of the expression pattern of a huge number of genes across different experimental conditions. Biclustering in microarray data is an important technique that discovers a group of genes that are coregulated in a subset of conditions. Biclustering algorithms require to identify coherent and nontrivial biclusters, i.e., the biclusters should have low mean squared residue and high row variance. A multiobjective genetic biclustering technique is proposed here that optimizes these objectives simultaneously. A novel encoding scheme that uses variable chromosome length is developed. Moreover, a new quantitative measure to evaluate the goodness of the biclusters is proposed. The performance of the proposed algorithm has been evaluated on both simulated and real-life gene expression datasets, and compared with some other well-known biclustering techniques.
Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay
IEEE Trans. Inf. Technol. Biomed.1
2008 Multiobjective fuzzy biclustering in microarray data: Method and a new performance measure
abstract
Objective of any biclustering algorithm in microarray data is to discover a subset of genes that are expressed similarly in a subset of conditions. The boundaries of biclusters usually overlap as genes and conditions may belong to different biclusters with different membership degrees. Hence the notion of fuzzy sets is useful for discovering such overlapping biclusters. In this article an attempt has been made to develop a multiobjective genetic algorithm based approach for probabilistic fuzzy biclustering that minimizes the residual and maximizes cluster size and expression profile variance. A novel variable string length encoding has been proposed in this regard that encodes multiple biclusters in a single string. Also a new performance measure that reflects how a bicluster is statistically distinguished from the background is proposed. Performance of the proposed algorithm has been compared with some well known biclustering algorithms.
Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Michael Q. Zhang, Xuegong Zhang
IEEE Congress on Evolutionary Computation1
2008 Combining multiobjective fuzzy clustering and probabilistic ANN classifier for unsupervised pattern classification: Application to satellite image segmentation
abstract
An important approach to unsupervised pixel classification in remote sensing satellite imagery is to use clustering in the spectral domain. In this article, a recently proposed multiobjective fuzzy clustering scheme has been combined with artificial neural networks (ANN) based probabilistic classifier to yield better performance. The multiobjective technique is first used to produce a set of non-dominated solutions. A part of these solutions having high confidence level are then used to train the ANN classifier. Finally the remaining solutions are classified using the trained classifier. The performance of this technique has been compared with that of some other well- known algorithms for two artificial data sets and a IRS satellite image of the city of Calcutta.
Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Ujjwal Maulik
IEEE Congress on Evolutionary Computation3
2008 A framework for an artificial immunity and speech based navigation for mobile robots
abstract
In recent years speech recognition technology and immunity based algorithms have made an impact in various areas and are deployed for a wide range of applications. This paper describes a learning process of a mobile robot which takes speech input as commands and performs some navigation task through a distinct man-machine interaction with the application of the learning based on the Artificial Immune System. For this purpose a 4-channel radio controlled Wheelbot and Microsof’s Speech SDK for speech recognition is employed. The speech recognition system is trained to recognize defined commands and the robot has been designed to navigate based on the instruction through the Speech Commands. The position of the obstacles are learnt and avoided by the help of immune algorithm.
Chingtham Tejbanta Singh, Ujjwal Maulik
IEEE Congress on Evolutionary Computation2
2008 Unsupervised Pixel Classification in Satellite Imagery: A Two-stage Fuzzy Clustering Approach
Anirban Mukhopadhyay 0001, Ujjwal Maulik
Fundam. Informaticae2
2008 A Simulated Annealing-Based Multiobjective Optimization Algorithm: AMOSA
abstract
This paper describes a simulated annealing based multiobjective optimization algorithm that incorporates the concept of archive in order to provide a set of tradeoff solutions for the problem under consideration. To determine the acceptance probability of a new solution vis-a-vis the current solution, an elaborate procedure is followed that takes into account the domination status of the new solution with the current solution, as well as those in the archive. A measure of the amount of domination between two solutions is also used for this purpose. A complexity analysis of the proposed algorithm is provided. An extensive comparative study of the proposed algorithm with two other existing and well-known multiobjective evolutionary algorithms (MOEAs) demonstrate the effectiveness of the former with respect to five existing performance measures, and several test problems of varying degrees of difficulty. In particular, the proposed algorithm is found to be significantly superior for many objective test problems (e.g., 4, 5, 10, and 15 objective problems), while recent studies have indicated that the Pareto ranking-based MOEAs perform poorly for such problems. In a part of the investigation, comparison of the real-coded version of the proposed algorithm is conducted with a very recent multiobjective simulated annealing algorithm, where the performance of the former is found to be generally superior to that of the latter.
Sanghamitra Bandyopadhyay, Sriparna Saha 0001, Ujjwal Maulik, Kalyanmoy Deb
IEEE Trans. Evol. Comput.3
2008 Gene Identification: Classical and Computational Intelligence Approaches
abstract
Automatic identification of genes has been an actively researched area of bioinformatics. Compared to earlier attempts for finding genes, the recent techniques are significantly more accurate and reliable. Many of the current gene-finding methods employ computational intelligence techniques that are known to be more robust when dealing with uncertainty and imprecision. In this paper, a detailed survey on the existing classical and computational intelligence based methods for gene identification is carried out. This includes a brief description of the classical and computational intelligence methods before discussing their applications to gene finding. In addition, a long list of available gene finders is compiled. For the convenience of the readers, the list is enhanced by mentioning their corresponding web sites and commenting on the general approach adopted. An extensive bibliography is provided. Finally, some limitations of the current approaches and future directions are discussed.
Sanghamitra Bandyopadhyay, Ujjwal Maulik, Debadyuti Roy
IEEE Trans. Syst. Man Cybern. Part C2
2008 Hierarchical Pattern Discovery in Graphs
abstract
In this correspondence, an efficient technique for detecting repeating patterns in a graph is described. For this purpose, the searching capability of evolutionary programming is utilized for discovering patterns that are often repeating in such structural data. The approach adopted in this correspondence is hierarchical in nature. Once a pattern is discovered in a particular level of the hierarchy, the graph is compressed using it, and the substructure discovery algorithm is repeated with the compressed graph. The proposed technique is useful for mining knowledge from databases that can be conveniently represented as graphs. The importance of such an endeavor can hardly be overemphasized, given that substantial portion of data that are generated and collected is either structural in nature or is composed of parts and relations between the parts, which can be naturally represented as graphs. A typical example can be the structure of protein as well as computer-aided design circuits that have a natural graphical representation.
Ujjwal Maulik
IEEE Trans. Syst. Man Cybern. Part C1
2007 Multiobjective approach to categorical data clustering
abstract
Categorical data clustering has been gaining significant attention from researchers since the last few years, because most of the real life data sets are categorical in nature. In contrast to numerical domain, no natural ordering can be found among the elements of a categorical domain. Hence no inherent distance measure, like the Euclidean distance, would work to compute the distance between two categorical objects. Most of the clustering algorithms designed for categorical data are based on optimizing a single objective function. However, a single objective function is often not applicable for different kinds of categorical data sets. Motivated by this fact, in this article, the categorical data clustering problem has been modeled as a multiobjective optimization problem. A popular multiobjective genetic algorithm has been used in this regard to optimize two objectives simultaneously, thus generating a set of non-dominated solutions. The performance of the proposed algorithm has been compared with that of different well known categorical data clustering algorithms and demonstrated for a variety of synthetic and real life categorical data sets. Also a statistical significance test has been performed to establish the superiority of the proposed algorithm.
Anirban Mukhopadhyay 0001, Ujjwal Maulik
IEEE Congress on Evolutionary Computation2
2007 An improved algorithm for clustering gene expression data
abstract
Abstract Motivation: Recent advancements in microarray technology allows simultaneous monitoring of the expression levels of a large number of genes over different time points. Clustering is an important tool for analyzing such microarray data, typical properties of which are its inherent uncertainty, noise and imprecision. In this article, a two-stage clustering algorithm, which employs a recently proposed variable string length genetic scheme and a multiobjective genetic clustering algorithm, is proposed. It is based on the novel concept of points having significant membership to multiple classes. An iterated version of the well-known Fuzzy C-Means is also utilized for clustering. Results: The significant superiority of the proposed two-stage clustering algorithm as compared to the average linkage method, Self Organizing Map (SOM) and a recently developed weighted Chinese restaurant-based clustering method (CRC), widely used methods for clustering gene expression data, is established on a variety of artificial and publicly available real life data sets. The biological relevance of the clustering solutions are also analyzed. Contact: [email protected] Supplementary information: The processed and normalized data sets, supplementary figures, tables and other related materials are available at http://d.1asphost.com/anirbanmukhopadhyay/simmts.html
Sanghamitra Bandyopadhyay, Anirban Mukhopadhyay 0001, Ujjwal Maulik
Bioinform.3
2007 Binary object extraction using bi-directional self-organizing neural network (BDSONN) architecture with fuzzy context sensitive thresholding
Siddhartha Bhattacharyya 0001, Paramartha Dutta, Ujjwal Maulik
Pattern Anal. Appl.3
2007 Multiobjective Genetic Clustering for Pixel Classification in Remote Sensing Imagery
abstract
An important approach for unsupervised landcover classification in remote sensing images is the clustering of pixels in the spectral domain into several fuzzy partitions. In this paper, a multiobjective optimization algorithm is utilized to tackle the problem of fuzzy partitioning where a number of fuzzy cluster validity indexes are simultaneously optimized. The resultant set of near-Pareto-optimal solutions contains a number of nondominated solutions, which the user can judge relatively and pick up the most promising one according to the problem requirements. Real-coded encoding of the cluster centers is used for this purpose. Results demonstrating the effectiveness of the proposed technique are provided for numeric remote sensing data described in terms of feature vectors. Different landcover regions in remote sensing imagery have also been classified using the proposed technique to establish its efficiency
Sanghamitra Bandyopadhyay, Ujjwal Maulik, Anirban Mukhopadhyay 0001
IEEE Trans. Geosci. Remote. Sens.2
2006 Clustering using Multi-objective Genetic Algorithm and its Application to Image Segmentation
abstract
This article presents a multiobjective fuzzy genetic clustering technique employing real coded encoding of cluster centers. Recent research has shown that clustering techniques that optimize a single objective may not provide satisfactory result because no single validity measure works well on different kinds of data sets. This fact has motivated us to develop a multiobjective fuzzy genetic clustering method that optimizes multiple validity measures simultaneously. User can chose any partitioning result from the resultant set of non dominated solutions according to the problem requirements. A number of artificial and real-life data sets have been clustered using the proposed fuzzy clustering method. Also the proposed algorithm has been applied for segmentation of a remote sensing image to show its effectiveness in pixel classification.
Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Ujjwal Maulik
SMC3
2006 Clustering distributed data streams in peer-to-peer environments
Sanghamitra Bandyopadhyay, Chris Giannella, Ujjwal Maulik, Hillol Kargupta, Kun Liu 0001, Souptik Datta
Inf. Sci.3
2005 A study of some fuzzy cluster validity indices, genetic clustering and application to pixel classification
Malay Kumar Pakhira, Sanghamitra Bandyopadhyay, Ujjwal Maulik
Fuzzy Sets Syst.3
2004 Validity index for crisp and fuzzy clusters
Malay Kumar Pakhira, Sanghamitra Bandyopadhyay, Ujjwal Maulik
Pattern Recognit.3
2003 Efficient BIST design for sequential machines using FiF-FoF values in machine states
abstract
This paper introduces a novel BIST-quality metric termed as the FiF -- FoF (Fan-in-Factor & Fan-out-Factor) defined on FSM-states. Based on the FiF -- FoF analysis, an efficient scheme is presented that ensures all state codes appear with uniform likelyhood at the present state (PS) lines during the test phase. This results in higher fault efficiency in a BIST structure. Experimental results on MCNC benchmarks show that the scheme improves fault efficiency of sequential circuits significantly, with marginal area overhead.
Samir Roy, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Subhadip Basu, Biplab K. Sikdar
ASP-DAC2
2003 A goal programming procedure for fuzzy multiobjective linear fractional programming problem
Bijay Baran Pal, Bhola Nath Moitra, Ujjwal Maulik
Fuzzy Sets Syst.3
2003 Fuzzy partitioning using a real-coded variable-length genetic algorithm for pixel classification
abstract
The problem of classifying an image into different homogeneous regions is viewed as the task of clustering the pixels in the intensity space. Real-coded variable string length genetic fuzzy clustering with automatic evolution of clusters is used for this purpose. The cluster centers are encoded in the chromosomes, and the Xie-Beni index is used as a measure of the validity of the corresponding partition. The effectiveness of the proposed technique is demonstrated for classifying different landcover regions in remote sensing imagery. Results are compared with those obtained using the well-known fuzzy C-means algorithm.
Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE Trans. Geosci. Remote. Sens.1
2002 An evolutionary technique based on K-Means algorithm for optimal clustering in RN
Sanghamitra Bandyopadhyay, Ujjwal Maulik
Inf. Sci.2
2002 Performance Evaluation of Some Clustering Algorithms and Validity Indices
abstract
In this article, we evaluate the performance of three clustering algorithms, hard K-Means, single linkage, and a simulated annealing (SA) based technique, in conjunction with four cluster validity indices, namely Davies-Bouldin index, Dunn's index, Calinski-Harabasz index, and a recently developed index I. Based on a relation between the index I and the Dunn's index, a lower bound of the value of the former is theoretically estimated in order to get unique hard K-partition when the data set has distinct substructures. The effectiveness of the different validity indices and clustering methods in automatically evolving the appropriate number of clusters is demonstrated experimentally for both artificial and real-life data sets with the number of clusters varying from two to ten. Once the appropriate number of clusters is determined, the SA-based clustering technique is used for proper partitioning of the data into the said number of clusters.
Ujjwal Maulik, Sanghamitra Bandyopadhyay
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Genetic clustering for automatic evolution of clusters and application to image classification
Sanghamitra Bandyopadhyay, Ujjwal Maulik
Pattern Recognit.2
2002 Efficient prototype reordering in nearest neighbor classification
Sanghamitra Bandyopadhyay, Ujjwal Maulik
Pattern Recognit.2
2001 Clustering Using Simulated Annealing with Probabilistic Redistribution
abstract
An efficient partitional clustering technique, called SAKM-clustering, that integrates the power of simulated annealing for obtaining minimum energy configuration, and the searching capability of K-means algorithm is proposed in this article. The clustering methodology is used to search for appropriate clusters in multidimensional feature space such that a similarity metric of the resulting clusters is optimized. Data points are redistributed among the clusters probabilistically, so that points that are farther away from the cluster center have higher probabilities of migrating to other clusters than those which are closer to it. The superiority of the SAKM-clustering algorithm over the widely used K-means algorithm is extensively demonstrated for artificial and real life data sets.
Sanghamitra Bandyopadhyay, Ujjwal Maulik, Malay Kumar Pakhira
Int. J. Pattern Recognit. Artif. Intell.2
2001 SAFE: An Efficient Feature Extraction Technique
Ujjwal Maulik, Sanghamitra Bandyopadhyay, John C. Trinder
Knowl. Inf. Syst.1
2001 Nonparametric genetic clustering: comparison of validity indices
abstract
A variable-string-length genetic algorithm (GA) is used for developing a novel nonparametric clustering technique when the number of clusters is not fixed a-priori. Chromosomes in the same population may now have different lengths since they encode different number of clusters. The crossover operator is redefined to tackle the concept of variable string length. A cluster validity index is used as a measure of the fitness of a chromosome. The performance of several cluster validity indices, namely the Davies-Bouldin (1979) index, Dunn's (1973) index, two of its generalized versions and a recently developed index, in appropriately partitioning a data set, are compared.
Sanghamitra Bandyopadhyay, Ujjwal Maulik
IEEE Trans. Syst. Man Cybern. Syst.2
2000 Fault tolerant permutation mapping in multistage interconnection network
Ujjwal Maulik, Sanghamitra Bandyopadhyay, Siddhartha Bhattacharyya 0001
J. Syst. Archit.1
2000 Genetic algorithm-based clustering technique
Ujjwal Maulik, Sanghamitra Bandyopadhyay
Pattern Recognit.1
1998 Incorporating Chromosome Differentaition in Genetic Algorithms
Sanghamitra Bandyopadhyay, Sankar K. Pal, Ujjwal Maulik
Inf. Sci.3