Rajat K. De

dblp:62/2543 · also Rajat Kumar De · DBLP profile ↗
← Back
44ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-6080-1131ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 TiMePReSt: Time and memory efficient pipeline parallel DNN training with removed staleness
Ankita Dutta, Nabendu Chaki, Rajat K. De
Future Gener. Comput. Syst.3
2026 The Forward-Cooperation-Backward (FCB) learning in a multi-encoding uni-decoding neural network architecture
Prasun Dutta, Koustab Ghosh, Rajat K. De
Neurocomputing3
2026 Integrating state-space modeling, parameter estimation, deep learning, and docking techniques in drug repurposing: a case study on COVID-19 cytokine storm
abstract
OBJECTIVE: This study addresses the significant challenges posed by emerging SARS-CoV-2 variants, particularly in developing diagnostics and therapeutics. Drug repurposing is investigated by identifying critical regulatory proteins impacted by the virus, providing rapid and effective therapeutic solutions for better disease management. MATERIALS AND METHODS: We employed a comprehensive approach combining mathematical modeling and efficient parameter estimation to study the transient responses of regulatory proteins in both normal and virus-infected cells. Proportional-integral-derivative (PID) controllers were used to pinpoint specific protein targets for therapeutic intervention. Additionally, advanced deep learning models and molecular docking techniques were applied to analyse drug-target and drug-drug interactions, ensuring both efficacy and safety of the proposed treatments. This approach was applied to a case study focused on the cytokine storm in COVID-19, centering on Angiotensin-converting enzyme 2 (ACE2), which plays a key role in SARS-CoV-2 infection. RESULTS: Our findings suggest that activating ACE2 presents a promising therapeutic strategy, whereas inhibiting AT1R seems less effective. Deep learning models, combined with molecular docking, identified Lomefloxacin and Fostamatinib as stable drugs with no significant thermodynamic interactions, suggesting their safe concurrent use in managing COVID-19-induced cytokine storms. DISCUSSION: The results highlight the potential of ACE2 activation in mitigating lung injury and severe inflammation caused by SARS-CoV-2. This integrated approach accelerates the identification of safe and effective treatment options for emerging viral variants. CONCLUSION: This framework provides an efficient method for identifying critical regulatory proteins and advancing drug repurposing, contributing to the rapid development of therapeutic strategies for COVID-19 and future global pandemics.
Abhisek Bakshi, Kaustav Gangopadhyay, Sujit Basak, Rajat K. De, Souvik Sengupta 0001, Abhijit Dasgupta 0001
J. Am. Medical Informatics Assoc.4
2026 G-NeuroDAVIS: A generative model for data visualization through a generalized embedding
abstract
Visualizing high-dimensional datasets through a generalized embedding has been a longstanding challenge. Several methods have been proposed for this purpose, but they have yet to generate a generalized embedding that not only reveals the hidden patterns present in the data but also generates realistic high-dimensional samples from it. Motivated by this aspect, in this article, a novel generative model called G-NeuroDAVIS has been developed, which is capable of visualizing high-dimensional data through a generalized embedding and thereby generating new samples. The model leverages advanced generative techniques to produce high-quality embedding that captures the underlying structure of the data more effectively compared with the existing methods. G-NeuroDAVIS can be trained in both supervised and unsupervised settings. We have rigorously evaluated our model through a series of experiments, demonstrating superior performance in several downstream tasks, which highlights the effectiveness of the learned representations. Results of an interpolation experiment reflect a smooth and meaningful transition in the generated images across various paths, which in turn depict preservation of underlying data structure. Furthermore, the conditional sample generation capability of the model has been described through both qualitative and quantitative assessments, revealing a marked improvement in generating realistic and diverse samples. G-NeuroDAVIS has outperformed Variational Autoencoder (VAE) significantly in terms of embedding quality and downstream tasks like classification. Moreover, the superior sample generation capability of G-NeuroDAVIS has been demonstrated against VAE, Deep Convolutional Generative Adversarial Network (DCGAN), Denoising Diffusion Probabilistic Models (DDPM), and Autoencoder (AE)-guided Real-valued Non-Volume Preserving (RealNVP). These results highlight the efficacy of G-NeuroDAVIS to serve as a robust tool in various applications that demand high-quality data generation and representation learning.
Chayan Maitra, Rajat K. De
Neural Networks2
2025 Efficient parameter estimation in biochemical pathways: Overcoming data limitations with constrained regularization and fuzzy inference
Abhisek Bakshi, Souvik Sengupta 0001, Rajat K. De, Abhijit Dasgupta 0001
Expert Syst. Appl.3
2024 NeuroDAVIS-FS: Feature Selection Through Visualization Using NeuroDAVIS
Chayan Maitra, Anwesha Sengupta, Rajat K. De
ICPR (26)3
2024 NeuroDAVIS: A neural network model for data visualization
Chayan Maitra, Dibyendu Bikash Seal, Rajat K. De
Neurocomputing3
2024 DN3MF: deep neural network for non-negative matrix factorization towards low rank approximation
Prasun Dutta, Rajat K. De
Pattern Anal. Appl.2
2024 ENLIGHTENMENT: A Scalable Annotated Database of Genomics and NGS-Based Nucleotide Level Profiles
abstract
The revolution in sequencing technologies has enabled human genomes to be sequenced at a very low cost and time leading to exponential growth in the availability of whole-genome sequences. However, the complete understanding of our genome and its association with cancer is a far way to go. Researchers are striving hard to detect new variants and find their association with diseases, which further gives rise to the need for aggregation of this Big Data into a common standard scalable platform. In this work, a database named Enlightenment has been implemented which makes the availability of genomic data integrated from eight public databases, and DNA sequencing profiles of H. sapiens in a single platform. Annotated results with respect to cancer specific biomarkers, pharmacogenetic biomarkers and its association with variability in drug response, and DNA profiles along with novel copy number variants are computed and stored, which are accessible through a web interface. In order to overcome the challenge of storage and processing of NGS technology-based whole-genome DNA sequences, Enlightenment has been extended and deployed to a flexible and horizontally scalable database HBase, which is distributed over a hadoop cluster, which would enable the integration of other omics data into the database for enlightening the path towards eradication of cancer.
Rituparna Sinha, Rajat Kumar Pal, Rajat K. De
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 UMINT-FS: UMINT-guided Feature Selection for multi-omics datasets
abstract
Feature selection is a crucial step in single-cell biological data analysis. It involves identifying and selecting a subset of features (genes, proteins, peaks among others) that are most informative and relevant for downstream analysis. A prior investigation has introduced an unsupervised neural network model, known as UMINT, tailored for the integration of single-cell multi-omics data. This novel deep learning model excels at single-cell multi-omics integration and feature extraction, yet lacks the ability to perform feature selection. The present study extends UMINT and introduces UMINT-FS that enables selection of top features from multi-omics datasets by analysing the weights learned by the UMINT network during integration of the omics modalities. UMINT-FS can operate in both supervised and unsupervised learning environments. A supervised learning environment empowers it to find cell-type-specific markers. The performance of UMINT-FS has been evaluated on two different types of single-cell multi-omics datasets and results demonstrated better performance than current state-of-the-art methods.
Chayan Maitra, Dibyendu Bikash Seal, Vivek Das, Yevgeniy Vorobeychik, Rajat K. De
BIBM5
2023 A Novel CRISPR-MultiTargeter Multi-agent Reinforcement learning (CMT-MARL) algorithm to identify editable target regions using a Hybrid scoring from multiple similar sequences
Susobhan Baidya, Sankhayan Choudhury, Rajat K. De
Appl. Intell.3
2023 CASSL: A cell-type annotation method for single cell transcriptomics data using semi-supervised learning
Dibyendu Bikash Seal, Vivek Das, Rajat K. De
Appl. Intell.3
2022 scARMF: Association Rule Mining-based feature selection Framework for Single-Cell transcriptomics data
abstract
Single-cell RNA-sequencing (scRNA-seq) technologies have allowed researchers to investigate transcriptional regulation at a cellular resolution. One such analysis often involves extracting statistically significant groups of cells identified as clusters that enable cell-type identification, based on the presence or absence of canonical markers. However, it has been observed that cells with similar gene expression profiles, m ay sometimes represent variable transcriptional states. Identifying cell-type specific markers, is hence, not sufficient enough to understand the underlying molecular activity within a particular cell cluster. Rather, we should focus on finding key regulators within cell clusters. In order to assess cells’ functionality beyond marker-based studies, genes driving or being driven by these key regulators need to be analysed against reference databases. In this work, we have developed an Association Rule Mining (ARM)-based feature selection Framework, called scARMF, which can identify major gene-gene interactions within a specific cell-cluster of interest in scRNA-seq data. These interaction networks have helped us identify key regulatory hubs (genes), some of which have been found to be relevant canonical markers when validated against a benchmarked reference database. The sub-networks formed by hub genes along with their neighbours, have been further assessed via Over Representation Analysis (ORA)-based pathway enrichment. This has revealed interesting functional characteristics that can be important for further downstream biological interpretations.
Dibyendu Bikash Seal, Vivek Das, Rajat K. De
BIBM3
2022 Block Search Stochastic Simulation Algorithm (BlSSSA): A Fast Stochastic Simulation Algorithm for Modeling Large Biochemical Networks
abstract
Stochastic simulation algorithms are extensively used for exploring stochastic behavior of biochemical pathways/networks. Computational cost of these algorithms is high in simulating real biochemical systems due to their large size, complex structure and stiffness. In order to reduce the computational cost, several algorithms have been developed. It is observed that these algorithms are basically fast in simulating weakly coupled networks. In case of strongly coupled networks, they become slow as their computational cost become high in maintaining complex data structures. Here, we develop Block Search Stochastic Simulation Algorithm (BlSSSA). BlSSSA is not only fast in simulating weakly coupled networks but also fast in simulating strongly coupled and stiff networks. We compare its performance with other existing algorithms using two hypothetical networks, viz., linear chain and colloidal aggregation network, and three real biochemical networks, viz., B cell receptor signaling network, FceRI signaling network and a stiff 1,3-Butadiene Oxidation network. It has been shown that BlSSSA is faster than other algorithms considered in this study.
Debraj Ghosh, Rajat K. De
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 GenSeg and MR-GenSeg: A Novel Segmentation Algorithm and its Parallel MapReduce Based Approach for Identifying Genomic Regions With Copy Number Variations
abstract
Identifying intragenic as well as intergenic sequences of the DNA, having structural alterations, is a significantly important research area, since this may be the root cause of many neurological and autoimmune diseases, including cancer. Working with whole genome NGS data has provided a new insight in this regard, but has lead to huge explosion of data that is growing exponentially. Hence, the challenges lie in efficient means of storage and processing this big data. In this study, we have developed a novel segmentation algorithm, called GenSeg, and its parallel MapReduce based algorithm, called MR-GenSeg, for detecting copy number variations. In order to annotate CNVs (variants), segments formed by GenSeg/MR-GenSeg have been represented in a novel way using a binary tree, where each node is a CNV event. GenSeg considers each position specific data of whole genome DNA sequence, so that precise identification of breakpoints is possible. GenSeg/MR-GenSeg has been compared with twelve popular CNV detection algorithms, where it has outperformed the others in terms of sensitivity, and has achieved a good F-score value. MR-GenSeg has excelled in terms of SpeedUp, when compared with these algorithms. The effect of CNVs on immunoglobulin (IG) genes has also been analysed in this study. Availability: The source codes are available at https://github.com/rituparna-sinha/MapReduce-GENSEG.
Rituparna Sinha, Rajat Kumar Pal, Rajat K. De
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Pattern and Rule Mining for Identifying Signatures of Epileptic Patients from Clinical EEG Data
abstract
Epilepsy is a neurological condition of human being, mostly treated based on the patients’ seizure symptoms, often recorded over multiple visits to a health-care facility. The lengthy time-consuming process of obtaining multiple recordings creates an obstacle in detecting epileptic patients in real time. An epileptic signature validated over EEG data of multiple similar kinds of epilepsy cases will haste the decision-making process of clinicians. In this paper, we have identified EEG data derived signatures for differentiating epileptic patients from normal individuals. Here we define the signatures with the help of various machine learning techniques, viz., feature selection and classification, pattern mining, and fuzzy rule mining. These signatures will add confidence to the decision-making process for detecting epileptic patients. Moreover, we define separate signatures by incorporating few demographic features like gender and age. Such signatures may aid the clinicians with the generalized epileptic signature in case of complex decisions.
Abhijit Dasgupta 0001, Losiana Nayak, Ritankar Das, Debasis Basu, Preetam Chandra, Rajat K. De
Fundam. Informaticae6
2020 ASAPP: Architectural Similarity-Based Automated Pathway Prediction System and Its Application in Host-Pathogen Interactions
abstract
The significance of metabolic pathway prediction is to envision the viable unknown transformations that can occur provided the appropriate enzymes are present. It can facilitate the prediction of the consequences of host-pathogen interactions. In this article, we have proposed a new algorithm Architectural Similarity-based Automated Pathway Prediction (ASAPP) to predict metabolic pathways based on the structural similarity among the metabolites. ASAPP takes two-dimensional structure and molecular weight of metabolites as input, and generates a list of probable transformations without the knowledge of any externally established reactions, with an accuracy of 85.09 percent. ASAPP has also been applied to predict the outcome of pathogen liberated toxins on the carbohydrate and lipid pathways of the hosts. We have analyzed the disruption of host pathways in the presence of toxins, and have found that some metabolites in Glycolysis and the TCA cycle have a high chance of being the breakpoints in the pathway. The tool is available at http://asapp.droppages.com/.
Rishika Sen, Somnath Tagore, Rajat K. De
IEEE ACM Trans. Comput. Biol. Bioinform.3
2019 A Model for Distributed Processing and Analyses of NGS Data under Map-Reduce Paradigm
abstract
Massively parallel sequencing technique, introduced by NGS technology, has resulted in an exponential growth of sequencing data, with greatly reduced cost and increased throughput. This huge explosion of data has introduced new challenges in regard to its storage, integration, processing, and analyses. In this paper, we have proposed a novel distributed model under Map-Reduce paradigm to address the NGS big data problem. The architecture of the model involves Map-Reduce based modularized approach involving three different phases that support various analytical pipelines. The first phase will generate detailed base level information of various individual genomes, by granulating the alignment data. The other two phases independently process this base level information in parallel. One of these two phases will provide an integrated DNA profile of multiple individuals, whereas the other phase will generate contigs with similar features in an individual. Each of these three phases will generate a repository of genomic information that will facilitate other analytical pipelines. A simulated and real experimental prototypes has been provided as results to show the effectiveness of the model and its superiority over a few existing popular models and tools. A detailed description of the scope of applications of this model is also included in this article.
Sandip Samaddar, Rituparna Sinha, Rajat K. De
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 Analyzing epileptogenic brain connectivity networks using clinical EEG data
abstract
Epileptogenic brain connectivity networks are altered compared to normal ones. Here, we have investigated the properties of epileptogenic networks by applying graph theoretical, statistical and machine learning approaches to the resting state electroencephalography (EEG) recordings obtained from 30 normal volunteers and 51 patients suffering from generalized epilepsy. In the case of epileptic patients, we have found that the brain networks behave like random networks. There is some loss in node connectivity. Hub nodes are more affected during epilepsy. Hence, the epileptogenic networks show less clustering coefficient than normal ones. In addition, we have identified 11 specific regions of brains and ten most significant connections among them as an epileptogenic signature by feature extraction. The ten most significant features are used to classify 81 sample data sets into two classes, i.e., epileptogenic and normal, with 79.01% accuracy. The highly probable eleven regions of human brain according to the positions of electrodes and connections among them may lead to a progress in the clinical treatment of epileptic patients.
Abhijit Dasgupta 0001, Ritankar Das, Losiana Nayak, Rajat K. De
BIBM4
2015 A module tree of Wnt signal transduction pathways
abstract
Wnt signal transduction pathway (Wnt STP) is a crucial intracellular pathway mainly due to its participation in important biological functions, i.e., embryonic development, and stem-cell management among others as well as in human pathology, mainly cancer. For these very reasons, Wnt STP is one of the highest researched signal transduction pathways. Study and analysis of its origin, expansion and gradual development to the present state as seen in human beings is one facet of this multi-pronged research. Development of the Wnt signal transduction pathway in cellular environment among various species is not clear till date. A phylogenetic tree obtained from Wnt STPs of multiple species is one of the ways to handle this problem. In this respect, we propose a new idea of constructing a phylogenetic tree from modules of Wnt STPs of diverse species. We term it as the `Module Tree'. A module is nothing but a self-sufficient minimally-dependent subset of the original Wnt STP. Authenticity of the module tree is tested by comparing it with the two reference trees. It performs better than an alternative phylogenetic tree constructed from pathway topology of Wnt STPs.
Losiana Nayak, Nitai P. Bhattacharyya, Rajat K. De
BIBM3
2015 A novel locally guided genome reassembling technique using an artificial ant system
Susobhan Baidya, Rajat K. De
Appl. Intell.2
2014 Selection of genes mediating certain cancers, using a neuro-fuzzy approach
Anupam Ghosh, Bibhas Chandra Dhara, Rajat K. De
Neurocomputing3
2013 An Optimization Rule for In Silico Identification of Targeted Overproduction in Metabolic Pathways
abstract
In an extension of previous work, here we introduce a second-order optimization method for determining optimal paths from the substrate to a target product of a metabolic network, through which the amount of the target is maximum. An objective function for the said purpose, along with certain linear constraints, is considered and minimized. The basis vectors spanning the null space of the stoichiometric matrix, depicting the metabolic network, are computed, and their convex combinations satisfying the constraints are considered as flux vectors. A set of other constraints, incorporating weighting coefficients corresponding to the enzymes in the pathway, are considered. These weighting coefficients appear in the objective function to be minimized. During minimization, the values of these weighting coefficients are estimated and learned. These values, on minimization, represent an optimal pathway, depicting optimal enzyme concentrations, leading to overproduction of the target. The results on various networks demonstrate the usefulness of the methodology in the domain of metabolic engineering. A comparison with the standard gradient descent and the extreme pathway analysis technique is also performed. Unlike the gradient descent method, the present method, being independent of the learning parameter, exhibits improved results.
Mouli Das, Late C. A. Murthy, Rajat K. De
IEEE ACM Trans. Comput. Biol. Bioinform.3
2011 A novel noise handling method to improve clustering of gene expression patterns
abstract
Cluster analysis of gene expression data is a useful tool for identifying biologically relevant groups of genes that show similar expression patterns under multiple experimental conditions. Performance of clustering algorithms is largely dependent on selected similarity measure. Efficiency in handling outliers is a major contributor to the success of a similarity measure. In gene expression data, there may be pairs of genes that have completely different expression values over a few samples under certain experimental condition(s), although they exhibit similar behavior over the other samples. Depending on the algorithms, these outliers are either placed in single element clusters (hierarchical clustering), are allowed to be in a cluster that is more similar compared to others (partitioning clustering) or they may be completely discarded from grouping (density-based, grid-based and graph-based clustering). In all these cases outliers affect the outcome of a clustering result. Measurement errors or conditional changes during microarray experiments may cause a single sample, if not more, differing in expression level to a great extent compared to the other samples. Expression value of the single or a very few outlier samples may cause a gene to be an outlier. We formulate a new weighted function based method to reduce the effect of outliers on similarity measures. The better the similarity measure is in measuring similarity between genes in the presence of outliers, the better the performance of the clustering algorithm will be in forming biologically relevant groups of genes. The effectiveness of the weighted function based method has been demonstrated with the clustering algorithms, viz. , K-means [ 1 ], Minimization of Disagreement (MIND) [ 2 ], Divisive Correlation Clustering Algorithm (DCCA) [ 3 ], Average Correlation Clustering Algorithm (ACCA) [ 4 ] and Bi-Correlation Clustering Algorithm (BCCA) [ 5 ] on a yeast gene expression dataset (Yeast Cheng and Church dataset from Yeast Functional Genomics Database [ http://yfgdb.princeton.edu/ ]). Assessment of the results has been done by using P-values on functional annotations. P-values less than 5.0 × 10 are reported as enriched functional categories. Figure 1 shows the number of functionally enriched attributes in the most enriched clusters obtained by each of the clustering and biclustering algorithms on the yeast gene expression dataset. The results suggest that the new weighted function based method significantly improves performance of all the cases, in terms of finding biologically relevant groups of genes. Number of functionally enriched attributes in the most enriched clusters obtained by different clustering and biclustering algorithms.
Anindya Bhattacharya, Rajat K. De
BMC Bioinform.2
2010 Average correlation clustering algorithm (ACCA) for grouping of co-regulated genes with similar pattern of variation in their expression values
Anindya Bhattacharya, Rajat K. De
J. Biomed. Informatics2
2009 Bi-correlation clustering algorithm for determining a set of co-regulated genes
abstract
MOTIVATION: Biclustering has been emerged as a powerful tool for identification of a group of co-expressed genes under a subset of experimental conditions (measurements) present in a gene expression dataset. Several biclustering algorithms have been proposed till date. In this article, we address some of the important shortcomings of these existing biclustering algorithms and propose a new correlation-based biclustering algorithm called bi-correlation clustering algorithm (BCCA). RESULTS: BCCA has been able to produce a diverse set of biclusters of co-regulated genes over a subset of samples where all the genes in a bicluster have a similar change of expression pattern over the subset of samples. Moreover, the genes in a bicluster have common transcription factor binding sites in the corresponding promoter sequences. The presence of common transcription factors binding sites, in the corresponding promoter sequences, is an evidence that a group of genes in a bicluster are co-regulated. Biclusters determined by BCCA also show highly enriched functional categories. Using different gene expression datasets, we demonstrate strength and superiority of BCCA over some existing biclustering algorithms. AVAILABILITY: The software for BCCA has been developed using C and Visual Basic languages, and can be executed on the Microsoft Windows platforms. The software may be downloaded as a zip file from http://www.isical.ac.in/ approximately rajat. Then it needs to be installed. Two word files (included in the zip file) need to be consulted before installation and execution of the software. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anindya Bhattacharya, Rajat K. De
Bioinform.2
2009 Interval based fuzzy systems for identification of important genes from microarray gene expression data: Application to carcinogenic development
Rajat K. De, Anupam Ghosh
J. Biomed. Informatics1
2008 Determination of optimal metabolic pathways through a new learning algorithm
abstract
In the present article, we introduce a new method for identification of metabolic pathways in constraint based models that consider enzyme and substrate concentrations. It generates data on reaction fluxes based on biomass conservation constraint and then a set of constraints is formulated incorporating weighting coefficients corresponding to concentration of enzymes catalyzing reactions in the pathway. Finally, the rate of yield of the target metabolite, starting with a given substrate, is maximized in order to identify an optimal pathway through these weighting coefficients. In an attempt to solve this problem, we have developed a learning technique that optimizes a given objective function to find the optimal pathways. Finally, we propose a modification of the Newton Raphson method and incorporate it to our proposed methodology, which yields more relevant results from the perspective of biology.
Late C. A. Murthy, Mouli Das, Rajat K. De, Subhasis Mukhopadhyay
ICPR3
2008 Divisive Correlation Clustering Algorithm (DCCA) for grouping of genes: detecting varying patterns in expression profiles
abstract
MOTIVATION: Cluster analysis (of gene-expression data) is a useful tool for identifying biologically relevant groups of genes that show similar expression patterns under multiple experimental conditions. Various methods have been proposed for clustering gene-expression data. However most of these algorithms have several shortcomings for gene-expression data clustering. In the present article, we focus on several shortcomings of conventional clustering algorithms and propose a new one that is able to produce better clustering solution than that produced by some others. RESULTS: We present the Divisive Correlation Clustering Algorithm (DCCA) that is suitable for finding a group of genes having similar pattern of variation in their expression values. To detect clusters with high correlation and biological significance, we use the correlation clustering concept introduced by Bansal et al. Our proposed algorithm DCCA produces a clustering solution without taking number of clusters to be created as an input. DCCA uses the correlation matrix in such a way that all genes in a cluster have highest average correlation with genes in that cluster. To test the performance of the DCCA, we have applied DCCA and some well-known conventional methods to an artificial dataset, and nine gene-expression datasets, and compared the performance of the algorithms. The clustering results of the DCCA are found to be more significantly relevant to the biological annotations than those of the other methods. All these facts show the superiority of the DCCA over some others for the clustering of gene-expression data. AVAILABILITY: The software has been developed using C and Visual Basic languages, and can be executed on the Microsoft Windows platforms. The software may be downloaded as a zip file from http://www.isical.ac.in/~rajat. Then it needs to be installed. Two word files (included in the zip file) need to be consulted before installation and execution of the software.
Anindya Bhattacharya, Rajat K. De
Bioinform.2
2007 An algorithm for modularization of MAPK and calcium signaling pathways: Comparative analysis among different species
Losiana Nayak, Rajat K. De
J. Biomed. Informatics2
2006 Identification of Over and Under Expressed Genes Mediating Allergic Asthma
Rajat K. De, Anindya Bhattacharya
IEA/AIE1
2006 Connectionist Modelling of Dynamics of Gene Expression and Reverse Engineering Gene Regulatory Networks
abstract
In this article we develop two connectionist models describing the dynamics of gene expression incorporating protein concentration. The models are based on the theoretical study of Goutsias and Kim [37]. We calculate the concentration of mRNAs and proteins at different time steps, and the concentrations of mRNAs and proteins are calculated as a function of step n. Here we consider concentration of mRNA in a cell at step as depending on the concentration of mRNA and proteins at step (n — 1) in that particular cell. Similarly the protein concentration in a cell at step n depends on the concentration of protein and mRNA at step (n — 1) in that particular cell. Here we develop two neural network models, and estimate the parameters using neural network model through learning. Finally, gene regulatory networks are determined as network parameters. The performance of the models have effectively been tested on a real life fruit fly time series gene expression data containing various stages of development of fruit fly.
Rajat K. De, Kasturi Biswas
IJCNN1
2005 A hardware pipeline for function optimization using genetic algorithms
abstract
Genetic Algorithms (GAs) are very commonly used as function optimizers, basically due to their search capability. A number of different serial and parallel versions of GA exist. In this paper, a pipelined version of the commonly used Genetic Algorithms and a corresponding hardware platform is described. The main idea of achieving pipelined execution of different operations of GA is to use a stochastic selection function which works with the fitness value of the candidate chromosome only. The modified algorithm is termed PLGA (Pipelined Genetic Algorithm). When executed in a CGA (Classical Genetic Algorithm) framework, the stochastic selection gives comparable performances with the roulette-wheel selection. In the pipelined hardware environment, PLGA will be much faster than the CGA. When executed on similar hardware platforms, PLGA may attain a maximum speedup of four over CGA. However, if CGA is executed in a uniprocessor system the speedup is much more. A comparison of PLGA against PGA (Parallel Genetic Algorithms) shows that PLGA may be even more effective than PGAs. A scheme for realizing the hardware pipeline is also presented. Since a general function evaluation unit is essential, a detailed description of one such unit is presented.
Malay Kumar Pakhira, Rajat K. De
GECCO2
2003 Extraction of Features Using M-Band Wavelet Packet Frame and Their Neuro-Fuzzy Evaluation for Multitexture Segmentation
abstract
In this paper, we propose a scheme for segmentation of multitexture images. The methodology involves extraction of texture features using an overcomplete wavelet decomposition scheme called discrete M-band wavelet packet frame (DMbWPF). This is followed by the selection of important features using a neuro-fuzzy algorithm under unsupervised learning. A computationally efficient search procedure is developed for finding the optimal basis based on some maximum criterion of textural measures derived from the statistical parameters for each of the subbands. The superior discriminating capability of the extracted features for segmentation of various texture images over those obtained by several existing methods is established.
Mausumi Acharyya, Rajat K. De, Malay Kumar Kundu
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Segmentation of remotely sensed images using wavelet features and their evaluation in soft computing framework
abstract
The present paper describes a feature extraction method based on M-band wavelet packet frames for segmenting remotely sensed images. These wavelet features are then evaluated and selected using an efficient neurofuzzy algorithm. Both the feature extraction and neurofuzzy feature evaluation methods are unsupervised, and they do not require the knowledge of the number and distribution of classes corresponding to various land covers in remotely sensed images. The effectiveness of the methodology is demonstrated on two four-band Indian Remote Sensing 1A satellite (IRS-1A) images containing five to six overlapping classes and a three-band SPOT image containing seven overlapping classes.
Mausumi Acharyya, Rajat K. De, Malay Kumar Kundu
IEEE Trans. Geosci. Remote. Sens.2
2002 Unsupervised feature extraction using neuro-fuzzy approach
Rajat K. De, Jayanta Basak, Sankar K. Pal
Fuzzy Sets Syst.1
2001 A connectionist model for selection of cases
Rajat K. De, Sankar K. Pal
Inf. Sci.1
2000 Unsupervised feature evaluation: a neuro-fuzzy approach
abstract
The present article demonstrates a way of formulating neuro-fuzzy approaches for both feature selection and extraction under unsupervised learning. A fuzzy feature evaluation index for a set of features is defined in terms of degree of similarity between two patterns in both the original and transformed feature spaces. A concept of flexible membership function incorporating weighted distance is introduced for computing membership values in the transformed space. Two new layered networks are designed. The tasks of membership computation and minimization of the evaluation index, through unsupervised learning process, are embedded into them without requiring the information on the number of clusters in the feature space. The network for feature selection results in an optimal order of individual importance of the features. The other one extracts a set of optimum transformed features, by projecting -dimensional original space directly to n'-dimensional (n' < n) transformed space, along with their relative importance. The superiority of the networks to some related ones is established experimentally.
Sankar K. Pal, Rajat K. De, Jayanta Basak
IEEE Trans. Neural Networks Learn. Syst.2
1999 Neuro-fuzzy feature evaluation with theoretical analysis
Rajat K. De, Jayanta Basak, Sankar K. Pal
Neural Networks1
1998 Fuzzy Feature Evaluation Index and Connectionist Realization - II. Theoretical Analysis
Jayanta Basak, Rajat K. De, Sankar K. Pal
Inf. Sci.2
1998 Fuzzy Feature Evaluation Index and Connectionist Realization
Sankar K. Pal, Jayanta Basak, Rajat K. De
Inf. Sci.3
1998 Unsupervised feature selection using a neuro-fuzzy approach
Jayanta Basak, Rajat K. De, Sankar K. Pal
Pattern Recognit. Lett.2
1997 Feature analysis: Neural network and fuzzy set theoretic approaches
Rajat K. De, Nikhil R. Pal, Sankar K. Pal
Pattern Recognit.1
1997 Knowledge-based fuzzy MLP for classification and rule generation
abstract
A new scheme of knowledge-based classification and rule generation using a fuzzy multilayer perceptron (MLP) is proposed. Knowledge collected from a data set is initially encoded among the connection weights in terms of class a priori probabilities. This encoding also includes incorporation of hidden nodes corresponding to both the pattern classes and their complementary regions. The network architecture, in terms of both links and nodes, is then refined during training. Node growing and link pruning are also resorted to. Rules are generated from the trained network using the input, output, and connection weights in order to justify any decision(s) reached. Negative rules corresponding to a pattern not belonging to a class can also be obtained. These are useful for inferencing in ambiguous cases. Results on real life and synthetic data demonstrate that the speed of learning and classification performance of the proposed scheme are better than that obtained with the fuzzy and conventional versions of the MLP (involving no initial knowledge encoding). Both convex and concave decision regions are considered in the process.
Sushmita Mitra, Rajat K. De, Sankar K. Pal
IEEE Trans. Neural Networks2