Indrajit Saha

dblp:39/3217 · DBLP profile ↗
← Back
26ranked-venue papers
13as first author
4since 2021 · last 2025
0000-0001-9513-9707ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 1 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Coalitions on the Fly in Cooperative Games
abstract
In this work, we examine a sequential setting of a cooperative game in which players arrive dynamically to form coalitions and complete tasks either together or individually, depending on the value created. Upon arrival, a new player as a decision maker faces two options: forming a new coalition or joining an existing one. We assume that players are greedy, i.e., they aim to maximize their rewards based on the information available at their arrival. The objective is to design an online value distribution policy that incentivizes players to form a coalition structure that maximizes social welfare. We focus on monotone and bounded cooperative games. Our main result establishes an upper bound of 3min/max on the competitive ratio for any irrevocable policy (i.e., one without redistribution), and proposes a policy that achieves a near-optimal competitive ratio of min{1/2, 3min/max}, where min and max denote the smallest and largest marginal contribution of any sub-coalition of players respectively. Finally, we also consider non-irrevocable policies, with alternative bounds only when the number of players is limited.
Yao Zhang 0011, Indrajit Saha, Zhaohong Sun 0001, Makoto Yokoo
ECAI2
2025 Weighted Envy-free Allocation with Subsidy
Haris Aziz 0001, Kei Kimura, Indrajit Saha, Zhaohong Sun 0001, Mashbat Suzuki, Makoto Yokoo
AAMAS4
2024 Subtraction games in more than one dimension
abstract
This paper concerns two-player alternating play combinatorial games (Conway 1976) in the normal-play convention, i.e. last move wins. Specifically, we study impartial vector subtraction games on tuples of nonnegative integers (Golomb 1966), with finite subtraction sets. In case of two move rulesets we find a complete solution, via a certain P -to- P principle (where P means that the previous player wins). Namely x ∈ P if and only if x + a + b ∈ P , where a and b are the two move options. Flammenkamp (1997) observed that, already in one dimension, rulesets with three moves can be hard to analyze, and still today his related conjecture remains open. Here, we solve instances of rulesets with three moves in two dimensions, and conjecture that they all have regular outcomes. Through several computer visualizations of outcomes of multi-move two-dimensional rulesets, we observe that they tend to partition the game board into periodic mosaics on very few regions/segments, which can depend on the number of moves in a ruleset. For example, we have found a five-move ruleset with an outcome segmentation into six semi-infinite slices. In this spirit, we develop a coloring automaton that generalizes the P -to- P principle. Given an initial set of colored positions, it quickly paints the P -positions in segments of the game board. Moreover, we prove that two-dimensional rulesets have row/column eventually periodic outcomes. We pose open problems on the generic hardness of two-dimensional rulesets; several regularity conjectures are provided, but we also conjecture that not all rulesets have regular outcomes.
Urban Larsson, Indrajit Saha, Makoto Yokoo
Theor. Comput. Sci.2
2021 Whole genome analysis of more than 10 000 SARS-CoV-2 virus unveils global genetic diversity and target region of NSP6
abstract
Whole genome analysis of SARS-CoV-2 is important to identify its genetic diversity. Moreover, accurate detection of SARS-CoV-2 is required for its correct diagnosis. To address these, first we have analysed publicly available 10 664 complete or near-complete SARS-CoV-2 genomes of 73 countries globally to find mutation points in the coding regions as substitution, deletion, insertion and single nucleotide polymorphism (SNP) globally and country wise. In this regard, multiple sequence alignment is performed in the presence of reference sequence from NCBI. Once the alignment is done, a consensus sequence is build to analyse each genomic sequence to identify the unique mutation points as substitutions, deletions, insertions and SNPs globally, thereby resulting in 7209, 11700, 119 and 53 such mutation points respectively. Second, in such categories, unique mutations for individual countries are determined with respect to other 72 countries. In case of India, unique 385, 867, 1 and 11 substitutions, deletions, insertions and SNPs are present in 566 SARS-CoV-2 genomes while 458, 1343, 8 and 52 mutation points in such categories are common with other countries. In majority (above 10%) of virus population, the most frequent and common mutation points between global excluding India and India are L37F, P323L, F506L, S507G, D614G and Q57H in NSP6, RdRp, Exon, Spike and ORF3a respectively. While for India, the other most frequent mutation points are T1198K, A97V, T315N and P13L in NSP3, RdRp, Spike and ORF8 respectively. These mutations are further visualised in protein structures and phylogenetic analysis has been done to show the diversity in virus genomes. Third, a web application is provided for searching mutation points globally and country wise. Finally, we have identified the potential conserved region as target that belongs to the coding region of ORF1ab, specifically to the NSP6 gene. Subsequently, we have provided the primers and probes using that conserved region so that it can be used for detecting SARS-CoV-2. Contact:[email protected] information: Supplementary data are available at http://www.nitttrkol.ac.in/indrajit/projects/COVID-Mutation-10K.
Indrajit Saha, Nimisha Ghosh, Ayan Pradhan, Debasree Maity, Kaushik Mitra
Briefings Bioinform.1
2019 Identification of Epigenetic Biomarkers with the use of Gene Expression and DNA Methylation for Breast Cancer Subtypes
abstract
Breast cancer is one of the most deadly cancers. It has four subtypes: Luminal A (LA), Luminal B (LB), HER2-enriched (HER2-E) and Basal-like (BL). For the cause of breast cancer subtypes, there are different genetic and epigenetic factors involved in its progression and susceptibility. Thus, the identification of genetic and/or epigenetic biomarkers can be helpful to understand the biological mechanisms better and to improve the diagnostic processes of this disease and its subtypes. Hence, this fact motivated us to investigate the epigenetic factor, such as DNA Methylation, with the integration of gene expression in order to find epigenetic biomarkers for breast cancer subtypes. In this regard, we have identified set of up and down regulated genes for each subtype using differential analysis. Thereafter, regression based feature ranking problem is formed in order to find the DNA Methylation site that is mostly responsible for the change in expression of a gene, which is considered as an epigenetic biomarker. A bagging integrated ensemble of decision trees is used for the same. The results of top ten up and down regulated genes and their corresponding most significant DNA Methylation sites are reported for breast cancer subtypes. Moreover, these genes are validated visually by means of survival and expression plots, showing TF-Gene-DNA Methylation interactions, Protein-Protein interaction network, KEGG pathway and GO enrichment analysis. The results show that top differentially expressed up and down regulated genes viz. MMP11, NUF2, EXO1, HJURP, HOXA4, SYNM, CAV1 and COL4A3BP in breast cancer subtypes may change their expression because of DNA Methylation sites viz. cg22418565, cg26029744, cg24741598, cg04550103, cg25952581, cg02109162, cg18498156 and cg04985097 respectively. The code, datasets and supplementary material are present online11http://www.nitttrkol.ac.in/indrajit/projects/epigenetic-mrna-breastcancer-subtypes/.
Indrajit Saha, Somnath Rakshit, Michal Wlasnowolski, Dariusz Plewczynski
TENCON1
2019 Improved Fuzzy Clustering using Ensemble based Differential Evolution for Remote Sensing Image
abstract
Identification of homogeneous regions in a satellite image is essentially the clustering of pixels in intensity space. Importantly remote sensing image like satellite images contain varieties of land cover types. Some of the covers are significantly large areas whereas some are relatively smaller regions. Therefore, automatically detecting such wide varying areas is a challenging task. Hence, this fact motivated us to propose an improved clustering technique viz. Ensemble based Differential Evolution for Fuzzy Clustering (EDEFC). For this purpose, very recently developed three variants of differential evolution (DE) are used in order to perform the clustering with different set of solutions. As a result, better clustering solution yields from the ensemble of DEs by exhaustive exploration of search space. The proposed EDEFC technique is applied on two numeric remote sensing datasets and Indian Remote Sensing (IRS) satellite image of Kolkata. The results of the EDEFC are shown quantitatively and visually by comparing with eight other clustering techniques. Moreover, the statistical significance test has also been performed in order to judge the superiority of EDEFC.
Jnanendra Prasad Sarkar, Indrajit Saha, Ujjwal Maulik
TENCON2
2019 Integrated Rough Fuzzy Clustering for Categorical data Analysis
Indrajit Saha, Jnanendra Prasad Sarkar, Ujjwal Maulik
Fuzzy Sets Syst.1
2019 Genome-wide analysis of multi-view data of miRNA-seq to identify miRNA biomarkers for stomach cancer
Namrata Pant, Somnath Rakshit, Sushmita Paul, Indrajit Saha
J. Biomed. Informatics4
2018 Deep Learning for Integrated Analysis of Breast Cancer Subtype Specific Multi-omics Data
abstract
Breast cancer is a deadly disease which commonly occurs all over the world and has been found to be the largest cause of cancer in females. Its detection is still a major challenge, both from a computational and biological point of views. Next Generation Sequencing (NGS) techniques have accelerated the mapping of human genomes rapidly. Involvement of advanced NGS techniques reveals that multiple genetic molecules are responsible for the cause of breast cancer and its subtypes. However, the high volume of data that is produced by the NGS techniques is difficult to study because of their high dimensionality and complexity. Thus, the integrated study of multi-omics data is one of the major challenges in medical science. This fact motivated us to study the NGS based high throughput expression data of miRNAs and mRNAs as well as Beta values of DNA Methylation of the corresponding mRNAs. In this regard, first, these datasets, together consisting of 33564 features of 305 patients in five classes viz. Luminal A, Luminal B, HER2-enriched, Basal-like and Control, are analysed in an integrated fashion using deep learning technique to classify the breast cancer subtypes properly. Second, the results of the deep learning technique are further analysed in order to identify the deeply connected features, i.e. either miRNA or mRNA or DNA Methylation, which are pivotal in the classification of breast cancer subtypes as well as play a crucial role in its formation. For this purpose, a deep learning technique, called stacked autoencoder is used to encode/transform the features into a low dimensional space, which is then fed to the five well known classifiers for classification. Moreover, the same encoded data is used to select the potential features after performing multiplication with the original data and Bonferroni correction on the p-values produced by the one-sample t-test. The results have been validated quantitatively and through biological significance analysis where oncogene TP53 and tumor suppression gene BRCA1 have been found. These genes are known to play a crucial role in breast cancer. The datasets, code and supplementary materials of this work are provided online at http://www.nitttrkol.ac.in/indrajit/projects/integrated-analysis-breastcancer-subtypes/.
Somnath Rakshit, Indrajit Saha, Subha Shankar Chakraborty, Dariusz Plewczynski
TENCON2
2018 Machine Learning for Object Labelling
abstract
Identification of objects from an image using machine learning is an emerging research topic. It is known to the scientific community that the labelling of objects by human intelligence in the current scenario of high volume data consumes a large amount of time and produces an error of 5% as observed in the case of ImageNet. In this regard, machine learning methods have become more popular and effective as they produce better results. However, the performance of different machine learning methods vary while identifying objects from images. To address this fact, an object labelling problem has been formulated by considering the images that contain objects of 10 different classes either in the form of solid, hollow and mixed. In order to label the objects in such three types of images, two different approaches are followed. First, a training dataset is created of size 240 samples where 24 samples are present in each of the 10 classes. Here each class contains equal number of solid and hollow objects. In the second approach, entropy is used to reduce the number of samples in the training dataset needed to achieve similar performance as obtained in the first approach. Thereafter, both the training datasets are applied separately to the seven well-known machine learning methods and one classical method to obtain the final results on three test images. The performance of the methods is demonstrated in terms of accuracy as well as by providing annotated objects in images. A software named ObLab2018 has also been developed to annotate the objects from images. The software and the datasets are provided online at http://www.nitttrkol.ac.in/indrajit/projects/ObLab2018/.
Indrajit Saha, Somnath Rakshit, Tanay Ghosh
TENCON1
2017 Semi-Supervised Learning with the Integration of Fuzzy Clustering and Artificial Neural Network
Indrajit Saha, Nivriti Debnath
HIS1
2017 Improving Modified Differential Evolution for Fuzzy Clustering
Jnanendra Prasad Sarkar, Indrajit Saha, Anasua Sarkar, Ujjwal Maulik
HIS2
2016 A new evolutionary microRNA marker selection using next-generation sequencing data
abstract
Next-generation sequencing allows high-throughput measurements of non-coding RNA expression levels in tissues. Analysis of microRNAs (miRNAs) is particularly effective in differentiation of cancerous tissue samples, based on patterns of their expression levels. The paper presents a wrapper feature selection approach based on t-Distributed Stochastic Neighbor Embedding (t-SNE), Covariance Matrix Adaptation Evolution Strategy (CMA-Es) and Support Vector Machine (SVM). The advantage of t-SNE is amplification of pairwise similarities by the means of t-Student neighborhood function. The attributes are embedded into 1-D space to reveal similarities between the features. Such information is used by CMA-ES through real-valued encoding in order to model pairwise relations between miRNAs with covariance matrices. Finally, the wrapper uses SVM to evaluate the objective, which expresses the tradeoff between classification quality and the desired number of features. The approach is tested on eight different cancer types from The Cancer Genome Atlas. It allows to find small sets of miRNAs to differentiate cancer types from a single tumor class to the normal one with high certainty.
Adrian Lancucki, Indrajit Saha, Shib Sankar Bhowmick, Ujjwal Maulik, Piotr Lipinski
CEC2
2015 A new evolutionary gene selection technique
abstract
Microarray technology allows to investigate gene expression levels by analyzing high dimensional datasets of few samples. Selection of discriminative, differentially expressed genes from such datasets is important to differentiate, prognose and understand the underlying biological processes. In this regard, the paper presents a new evolutionary gene selection method based on Student-t Stochastic Neighbor Embedding (t-SNE), Differential Evolution (DE) and Support Vector Machine (SVM). Here the underlying classification task of SVM is used as an optimization problem of DE, while t-SNE provides better ordering of genes for selection purpose. Generally, t-SNE is used to reorder the genes in such a way so that similar genes are grouped together and dissimilar genes are kept further apart. These reordered genes are then fragmented into fixed-length partitions. Thereafter, from each partition, a gene is selected randomly to encode the initial population of DE along with the combination of its weight and threshold values in order to participate in fitness computation. In the final generation of DE, a subset of genes is selected based on higher classification accuracy. The proposed technique is tested on six publicly available microarray datasets concerning various cancerous tissues of Homo sapiens and yields a potential set of genes by providing prefect or nearly perfect classification accuracy. Moreover, the superiority of the proposed technique has been demonstrated in comparison with other widely used techniques. Finally, the achieved results have also been justified by a statistical test and allowed us to draw biological conclusions through the identification of Gene Ontologies.
Adrian Lancucki, Indrajit Saha, Piotr Lipinski
CEC2
2015 Ensemble based rough fuzzy clustering for categorical data
Indrajit Saha, Jnanendra Prasad Sarkar, Ujjwal Maulik
Knowl. Based Syst.1
2014 Incremental learning based multiobjective fuzzy clustering for categorical data
Indrajit Saha, Ujjwal Maulik
Inf. Sci.1
2014 Multi-level thresholding using quantum inspired meta-heuristics
Sandip Dey, Indrajit Saha, Siddhartha Bhattacharyya 0001, Ujjwal Maulik
Knowl. Based Syst.2
2012 SVMeFC: SVM Ensemble Fuzzy Clustering for Satellite Image Segmentation
abstract
The problem of unsupervised image segmentation of a satellite image in a number of homogeneous regions can be viewed as the task of clustering the pixels in the intensity space. This letter presents an approach that exploits the capability of some recently proposed fuzzy clustering techniques, as well as support vector machine (SVM) classifiers, to yield improved solutions. All the fuzzy clustering techniques are first used to produce a set of different clustering solutions. Each such solution has been improved by a novel technique based on an SVM classifier. Thereafter, the cluster-based similarity partition algorithm is used to create the final clustering solution from all improved ensemble solutions. Results demonstrating the effectiveness of the proposed technique are provided for numeric remote sensing data described in terms of feature vectors. Moreover, a remotely sensed image of Calcutta City has been segmented using the proposed technique to establish its utility. In addition, the additional information of this letter is given as supplementary at http://sysbio.icm.edu.pl/indra/SVMeFC.html.
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
IEEE Geosci. Remote. Sens. Lett.1
2011 PMAFC: A New Probabilistic Memetic Algorithm Based Fuzzy Clustering
Indrajit Saha, Ujjwal Maulik, Dariusz Plewczynski
ISMIS1
2011 Improvement of new automatic differential fuzzy clustering using SVM classifier for microarray analysis
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
Expert Syst. Appl.1
2011 Unsupervised and Supervised Learning Approaches Together for Microarray Analysis
abstract
In this article, a novel concept is introduced by using both unsupervised and supervised learning. For unsupervised learning, the problem of fuzzy clustering in microarray data as a multiobjective optimization is used, which simultaneously optimizes two internal fuzzy cluster validity indices to yield a set of Pareto-optimal clustering solutions. In this regards, a new multiobjective differential evolution based fuzzy clustering technique has been proposed. Subsequently, for supervised learning, a fuzzy majority voting scheme along with support vector machine is used to integrate the clustering information from all the solutions in the resultant Pareto-optimal set. The performances of the proposed clustering techniques have been demonstrated on five publicly available benchmark microarray data sets. A detail comparison has been carried out with multiobjective genetic algorithm based fuzzy clustering, multiobjective differential evolution based fuzzy clustering, single objective versions of differential evolution and genetic algorithm based fuzzy clustering as well as well known fuzzy c-means algorithm. While using support vector machine, comparative studies of the use of four different kernel functions are also reported. Statistical significance test has been done to establish the statistical superiority of the proposed multiobjective clustering approach. Finally, biological significance test has been carried out using a web based gene annotation tool to show that the proposed integrated technique is able to produce biologically relevant clusters of coexpressed genes.
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
Fundam. Informaticae1
2010 Real-coded differential crisp clustering for MRI brain image segmentation
abstract
In this paper, a segmentation technique of multi-spectral magnetic resonance image of the brain using a new differential evolution based crisp clustering is proposed. Real-coded encoding of the cluster centres is used for this purpose. Here assignments of points to different clusters are made based on the Euclidean distance. The proposed method is applied on several simulated T1-weighted, T2-weighted and proton density for normal and MS lesion magnetic resonance brain images. Superiority of the proposed method over genetic algorithm based crisp clustering, simulated annealing based crisp clustering, K-means and average linkage are demonstrated quantitatively. Segmentation obtained by differential evolution based crisp clustering technique is also compared with the available ground truth information. Also statistical analysis has been conducted to judge the effectiveness. Matlab version of the software is available at http://bio.icm.edu.pl/~darman/MRI.
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski
IEEE Congress on Evolutionary Computation1
2010 Use of Multiobjective Differential Fuzzy Clustering with ANN Classifier for Unsupervised Pattern Classification: Application to Microarray Analysis
Indrajit Saha
RCIS1
2010 Automatic Fuzzy Clustering Using Modified Differential Evolution for Image Classification
abstract
The problem of classifying an image into different homogeneous regions is viewed as the task of clustering the pixels in the intensity space. In particular, satellite images contain landcover types, some of which cover significantly large areas while some (e.g., bridges and roads) occupy relatively much smaller regions. Automatically detecting regions or clusters of such widely varying sizes is a challenging task. In this paper, a new real-coded modified differential evolution based automatic fuzzy clustering algorithm is proposed which automatically evolves the number of clusters as well as the proper partitioning from a data set. Here, the assignment of points to different clusters is done based on a Xie-Beni index where the Euclidean distance is taken into consideration. The effectiveness of the proposed technique is first demonstrated for two numeric remote sensing data described in terms of feature vectors and then in identifying different landcover regions in remote sensing imagery. The superiority of the new method is demonstrated by comparing it with other existing techniques like automatic clustering using improved differential evolution, classical differential evolution based automatic fuzzy clustering, variable length genetic algorithm based fuzzy clustering, and well known fuzzy C-means algorithm both qualitatively and quantitatively.
Ujjwal Maulik, Indrajit Saha
IEEE Trans. Geosci. Remote. Sens.2
2010 Integrating Clustering and Supervised Learning for Categorical Data Analysis
abstract
The problem of fuzzy clustering of categorical data, where no natural ordering among the elements of a categorical attribute domain can be found, is an important problem in exploratory data analysis. As a result, a few clustering algorithms with focus on categorical data have been proposed. In this paper, a modified differential evolution (DE)-based fuzzy c-medoids (FCMdd) clustering of categorical data has been proposed. The algorithm combines both local as well as global information with adaptive weighting. The performance of the proposed method has been compared with those using genetic algorithm, simulated annealing, and the classical DE technique, besides the FCMdd, fuzzy k-modes, and average linkage hierarchical clustering algorithm for four artificial and four real life categorical data sets. Statistical test has been carried out to establish the statistical significance of the proposed method. To improve the result further, the clustering method is integrated with a support vector machine (SVM), a well-known technique for supervised learning. A fraction of the data points selected from different clusters based on their proximity to the respective medoids is used for training the SVM. The clustering assignments of the remaining points are thereafter determined using the trained classifier. The superiority of the integrated clustering and supervised learning approach has been demonstrated.
Ujjwal Maulik, Sanghamitra Bandyopadhyay, Indrajit Saha
IEEE Trans. Syst. Man Cybern. Part A3
2009 Modified differential evolution based fuzzy clustering for pixel classification in remote sensing imagery
Ujjwal Maulik, Indrajit Saha
Pattern Recognit.2