Sumanta Ray

dblp:136/7802 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-3371-8516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Enhancing Single-Cell RNA-Seq Data Completeness With a Graph Learning Framework
abstract
Single cell RNA sequencing (scRNA-seq) is a powerful tool to capture gene expression snapshots in individual cells. However, a low amount of RNA in the individual cells results in dropout events, which introduce huge zero counts in the single cell expression matrix. We have developed VAImpute, a variational graph autoencoder based imputation technique that learns the inherent distribution of a large network/graph constructed from the scRNA-seq data leveraging copula correlation ($Ccor$) among cells/genes. The trained model is utilized to predict the dropouts events by computing the probability of all non-edges (cell-gene) in the network. We devise an algorithm to impute the missing expression values of the detected dropouts. The performance of the proposed model is assessed on both simulated and real scRNA-seq datasets, comparing it to established single-cell imputation methods. VAImpute yields significant improvements to detect dropouts, thereby achieving superior performance in cell clustering, detecting rare cells, and differential expression.
Snehalika Lall, Sumanta Ray, Sanghamitra Bandyopadhyay
IEEE Trans. Comput. Biol. Bioinform.2
2022 Deep variational graph autoencoders for novel host-directed therapy options against COVID-19
Sumanta Ray, Snehalika Lall, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Alexander Schönhuth
Artif. Intell. Medicine1
2022 sc-REnF: An entropy guided robust feature selection for single-cell RNA-seq data
abstract
Annotation of cells in single-cell clustering requires a homogeneous grouping of cell populations. Since single-cell data are susceptible to technical noise, the quality of genes selected prior to clustering is of crucial importance in the preliminary steps of downstream analysis. Therefore, interest in robust gene selection has gained considerable attention in recent years. We introduce sc-REnF [robust entropy based feature (gene) selection method], aiming to leverage the advantages of $R{\prime}{e}nyi$ and $Tsallis$ entropies in gene selection for single cell clustering. Experiments demonstrate that with tuned parameter ($q$), $R{\prime}{e}nyi$ and $Tsallis$ entropies select genes that improved the clustering results significantly, over the other competing methods. sc-REnF can capture relevancy and redundancy among the features of noisy data extremely well due to its robust objective function. Moreover, the selected features/genes can able to determine the unknown cells with a high accuracy. Finally, sc-REnF yields good clustering performance in small sample, large feature scRNA-seq data. Availability: The sc-REnF is available at https://github.com/Snehalikalall/sc-REnF.
Snehalika Lall, Abhik Ghosh, Sumanta Ray, Sanghamitra Bandyopadhyay
Briefings Bioinform.3
2022 A copula based topology preserving graph convolution network for clustering of single-cell RNA-seq data
abstract
Annotation of cells in single-cell clustering requires a homogeneous grouping of cell populations. There are various issues in single cell sequencing that effect homogeneous grouping (clustering) of cells, such as small amount of starting RNA, limited per-cell sequenced reads, cell-to-cell variability due to cell-cycle, cellular morphology, and variable reagent concentrations. Moreover, single cell data is susceptible to technical noise, which affects the quality of genes (or features) selected/extracted prior to clustering. Here we introduce sc-CGconv (copula based graph convolution network for single clustering), a stepwise robust unsupervised feature extraction and clustering approach that formulates and aggregates cell-cell relationships using copula correlation (Ccor), followed by a graph convolution network based clustering approach. sc-CGconv formulates a cell-cell graph using Ccor that is learned by a graph-based artificial intelligence model, graph convolution network. The learned representation (low dimensional embedding) is utilized for cell clustering. sc-CGconv features the following advantages. a. sc-CGconv works with substantially smaller sample sizes to identify homogeneous clusters. b. sc-CGconv can model the expression co-variability of a large number of genes, thereby outperforming state-of-the-art gene selection/extraction methods for clustering. c. sc-CGconv preserves the cell-to-cell variability within the selected gene set by constructing a cell-cell graph through copula correlation measure. d. sc-CGconv provides a topology-preserving embedding of cells in low dimensional space.
Snehalika Lall, Sumanta Ray, Sanghamitra Bandyopadhyay
PLoS Comput. Biol.2
2021 RgCop-A regularized copula based method for gene selection in single-cell RNA-seq data
abstract
Gene selection in unannotated large single cell RNA sequencing (scRNA-seq) data is important and crucial step in the preliminary step of downstream analysis. The existing approaches are primarily based on high variation (highly variable genes) or significant high expression (highly expressed genes) failed to provide stable and predictive feature set due to technical noise present in the data. Here, we propose RgCop, a novel regularized copula based method for gene selection from large single cell RNA-seq data. RgCop utilizes copula correlation (Ccor), a robust equitable dependence measure that captures multivariate dependency among a set of genes in single cell expression data. We formulate an objective function by adding l1 regularization term with Ccor to penalizes the redundant co-efficient of features/genes, resulting non-redundant effective features/genes set. Results show a significant improvement in the clustering/classification performance of real life scRNA-seq data over the other state-of-the-art. RgCop performs extremely well in capturing dependence among the features of noisy data due to the scale invariant property of copula, thereby improving the stability of the method. Moreover, the differentially expressed (DE) genes identified from the clusters of scRNA-seq data are found to provide an accurate annotation of cells. Finally, the features/genes obtained from RgCop is able to annotate the unknown cells with high accuracy.
Snehalika Lall, Sumanta Ray, Sanghamitra Bandyopadhyay
PLoS Comput. Biol.2
2021 Colored Network Motif Analysis by Dynamic Programming Approach: An Application in Host Pathogen Interaction Network
abstract
Network motifs are subgraphs of a network which are found with significantly higher frequency than that expected in similar random networks. Motifs are small building blocks of a network and they have emerged as a way to uncover topological properties of complex networks. A special yet not much explored type of motif is the 'colored motif' where color (type) of each node, and hence the edges, in the motif is distinguishable from each other. A traditional motif is defined as a recurring structure in a network, whereas colored motif introduces detailed information about the color of the nodes. G-trie is a data structure to efficiently store a given set of subgraphs by exploiting the topological overlaps within them. In this article we have implemented a modified g-trie to store colored subgraphs and developed a method to discover colored motifs. Our method uses an approximate enumeration for counting the subgraphs to reduce the runtime. We have applied our method to find colored motifs of size three in a host pathogen protein-protein interaction network having two types of proteins namely HIV-1 and human proteins, and four types of edges. Here, we have discovered eight motifs, six of which contain both HIV-1 and human proteins, while the remaining two contain only human proteins.
Sumanta Ray, Sanghamitra Bandyopadhyay
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 Discovering Perturbation of Modular Structure in HIV Progression by Integrating Multiple Data Sources Through Non-Negative Matrix Factorization
abstract
Detecting perturbation in modular structure during HIV-1 disease progression is an important step to understand stage specific infection pattern of HIV-1 virus in human cell. In this article, we proposed a novel methodology on integration of multiple biological information to identify such disruption in human gene module during different stages of HIV-1 infection. We integrate three different biological information: gene expression information, protein-protein interaction information, and gene ontology information in single gene meta-module, through non negative matrix factorization (NMF). As the identified meta-modules inherit those information so, detecting perturbation of these, reflects the changes in expression pattern, in PPI structure and in functional similarity of genes during the infection progression. To integrate modules of different data sources into strong meta-modules, NMF based clustering is utilized here. Perturbation in meta-modular structure is identified by investigating the topological and intramodular properties and putting rank to those meta-modules using a rank aggregation algorithm. We have also analyzed the preservation structure of significant GO terms in which the human proteins of the meta-modules participate. Moreover, we have performed an analysis to show the change of coregulation pattern of identified transcription factors (TFs) over the HIV progression stages.
Sumanta Ray, Ujjwal Maulik
IEEE ACM Trans. Comput. Biol. Bioinform.1
2017 Preservation affinity in consensus modules among stages of HIV-1 progression
abstract
BACKGROUND: Analysis of gene expression data provides valuable insights into disease mechanism. Investigating relationship among co-expression modules of different stages is a meaningful tool to understand the way in which a disease progresses. Identifying topological preservation of modular structure also contributes to that understanding. METHODS: HIV-1 disease provides a well-documented progression pattern through three stages of infection: acute, chronic and non-progressor. In this article, we have developed a novel framework to describe the relationship among the consensus (or shared) co-expression modules for each pair of HIV-1 infection stages. The consensus modules are identified to assess the preservation of network properties. We have investigated the preservation patterns of co-expression networks during HIV-1 disease progression through an eigengene-based approach. RESULTS: We discovered that the expression patterns of consensus modules have a strong preservation during the transitions of three infection stages. In particular, it is noticed that between acute and non-progressor stages the preservation is slightly more than the other pair of stages. Moreover, we have constructed eigengene networks for the identified consensus modules and observed the preservation structure among them. Some consensus modules are marked as preserved in two pairs of stages and are analyzed further to form a higher order meta-network consisting of a group of preserved modules. Additionally, we observed that module membership (MM) values of genes within a module are consistent with the preservation characteristics. The MM values of genes within a pair of preserved modules show strong correlation patterns across two infection stages. CONCLUSIONS: We have performed an extensive analysis to discover preservation pattern of co-expression network constructed from microarray gene expression data of three different HIV-1 progression stages. The preservation pattern is investigated through identification of consensus modules in each pair of infection stages. It is observed that the preservation of the expression pattern of consensus modules remains more prominent during the transition of infection from acute stage to non-progressor stage. Additionally, we observed that the module membership values of genes are coherent with preserved modules across the HIV-1 progression stages.
Sk Md Mosaddek Hossain, Sumanta Ray, Anirban Mukhopadhyay 0001
BMC Bioinform.2
2017 A comprehensive analysis on preservation patterns of gene co-expression networks during Alzheimer's disease progression
abstract
BACKGROUND: Alzheimer's disease (AD) is a chronic neuro-degenerative disruption of the brain which involves in large scale transcriptomic variation. The disease does not impact every regions of the brain at the same time, instead it progresses slowly involving somewhat sequential interaction with different regions. Analysis of the expression patterns of the genes in different regions of the brain influenced in AD surely contribute for a enhanced comprehension of AD pathogenesis and shed light on the early characterization of the disease. RESULTS: Here, we have proposed a framework to identify perturbation and preservation characteristics of gene expression patterns across six distinct regions of the brain ("EC", "HIP", "PC", "MTG", "SFG", and "VCX") affected in AD. Co-expression modules were discovered considering a couple of regions at once. These are then analyzed to know the preservation and perturbation characteristics. Different module preservation statistics and a rank aggregation mechanism have been adopted to detect the changes of expression patterns across brain regions. Gene ontology (GO) and pathway based analysis were also carried out to know the biological meaning of preserved and perturbed modules. CONCLUSIONS: In this article, we have extensively studied the preservation patterns of co-expressed modules in six distinct brain regions affected in AD. Some modules are emerged as the most preserved while some others are detected as perturbed between a pair of brain regions. Further investigation on the topological properties of preserved and non-preserved modules reveals a substantial association amongst "betweenness centrality" and "degree" of the involved genes. Our findings may render a deeper realization of the preservation characteristics of gene expression patterns in discrete brain regions affected by AD.
Sumanta Ray, Sk Md Mosaddek Hossain, Lutfunnesa Khatun, Anirban Mukhopadhyay 0001
BMC Bioinform.1
2016 A NMF based approach for integrating multiple data sources to predict HIV-1-human PPIs
abstract
BACKGROUND: Predicting novel interactions between HIV-1 and human proteins contributes most promising area in HIV research. Prediction is generally guided by some classification and inference based methods using single biological source of information. RESULTS: In this article we have proposed a novel framework to predict protein-protein interactions (PPIs) between HIV-1 and human proteins by integrating multiple biological sources of information through non negative matrix factorization (NMF). For this purpose, the multiple data sets are converted to biological networks, which are then utilized to predict modules. These modules are subsequently combined into meta-modules by using NMF based clustering method. The integrated meta-modules are used to predict novel interactions between HIV-1 and human proteins. We have analyzed the significant GO terms and KEGG pathways in which the human proteins of the meta-modules participate. Moreover, the topological properties of human proteins involved in the meta modules are investigated. We have also performed statistical significance test to evaluate the predictions. CONCLUSIONS: Here, we propose a novel approach based on integration of different biological data sources, for predicting PPIs between HIV-1 and human proteins. Here, the integration is achieved through non negative matrix factorization (NMF) technique. Most of the predicted interactions are found to be well supported by the existing literature in PUBMED. Moreover, human proteins in the predicted set emerge as 'hubs' and 'bottlenecks' in the analysis. Low p-value in the significance test also suggests that the predictions are statistically significant.
Sumanta Ray, Sanghamitra Bandyopadhyay
BMC Bioinform.1
2016 Discovering Condition Specific Topological Pattern Changes in Coexpression Network: An Application to HIV-1 Progression
abstract
The natural progression of HIV-1 begins with a short acute retroviral syndrome which typically transit to chronic and clinical latency stages and subsequently progresses to a symptomatic, life-threatening immunodeficiency disease known as AIDS. Microarray analysis based on gene coexpression is widely used to investigate the coregulation pattern of a group (or cluster) of genes in a specific phenotype. Moreover, an investigation on the topological patterns across multiple phenotypes can facilitate the understanding of stage specific infection pattern of HIV-1 virus. Here, we develop a novel framework to identify topological patterns of gene co-expression network and detect changes of modular structure across different stages of HIV progression. This is achieved by comparing the topological and intramodular properties of HIV infection modules. To capture the diversity in modular structure, some topological, correlation based, and eigengene based measures are utilized here. We have applied a rank aggregation scheme to rank all the modules to provide a good agreement between these measures. Some novel transcription factors like 'FOXO1', 'GATA3', 'GFI1', 'IRF1', 'IRF7', 'MAX', 'STAT1', 'STAT3', 'XBP1', and 'YY1' that merge from the modules show significant change in expression pattern over HIV progression stages. Moreover, we have performed an eigengene based analysis to reveal the perturbation in modular structure across three stages of HIV-1 progression.
Sumanta Ray, Sanghamitra Bandyopadhyay
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 A review of in silico approaches for analysis and prediction of HIV-1-human protein-protein interactions
abstract
The computational or in silico approaches for analysing the HIV-1-human protein-protein interaction (PPI) network, predicting different host cellular factors and PPIs and discovering several pathways are gaining popularity in the field of HIV research. Although there exist quite a few studies in this regard, no previous effort has been made to review these works in a comprehensive manner. Here we review the computational approaches that are devoted to the analysis and prediction of HIV-1-human PPIs. We have broadly categorized these studies into two fields: computational analysis of HIV-1-human PPI network and prediction of novel PPIs. We have also presented a comparative assessment of these studies and proposed some methodologies for discussing the implication of their results. We have also reviewed different computational techniques for predicting HIV-1-human PPIs and provided a comparative study of their applicability. We believe that our effort will provide helpful insights to the HIV research community.
Sanghamitra Bandyopadhyay, Sumanta Ray, Anirban Mukhopadhyay 0001, Ujjwal Maulik
Briefings Bioinform.2
2014 Incorporating the type and direction information in predicting novel regulatory interactions between HIV-1 and human proteins using a biclustering approach
abstract
BACKGROUND: Discovering novel interactions between HIV-1 and human proteins would greatly contribute to different areas of HIV research. Identification of such interactions leads to a greater insight into drug target prediction. Some recent studies have been conducted for computational prediction of new interactions based on the experimentally validated information stored in a HIV-1-human protein-protein interaction database. However, these techniques do not predict any regulatory mechanism between HIV-1 and human proteins by considering interaction types and direction of regulation of interactions. RESULTS: Here we present an association rule mining technique based on biclustering for discovering a set of rules among human and HIV-1 proteins using the publicly available HIV-1-human PPI database. These rules are subsequently utilized to predict some novel interactions among HIV-1 and human proteins. For prediction purpose both the interaction types and direction of regulation of interactions, (i.e., virus-to-host or host-to-virus) are considered here to provide important additional information about the regulation pattern of interactions. We have also studied the biclusters and analyzed the significant GO terms and KEGG pathways in which the human proteins of the biclusters participate. Moreover the predicted rules have also been analyzed to discover regulatory relationship between some human proteins in course of HIV-1 infection. Some experimental evidences of our predicted interactions have been found by searching the recent literatures in PUBMED. We have also highlighted some human proteins that are likely to act against the HIV-1 attack. CONCLUSIONS: We pose the problem of identifying new regulatory interactions between HIV-1 and human proteins based on the existing PPI database as an association rule mining problem based on biclustering algorithm. We discover some novel regulatory interactions between HIV-1 and human proteins. Significant number of predicted interactions has been found to be supported by recent literature.
Anirban Mukhopadhyay 0001, Sumanta Ray, Ujjwal Maulik
BMC Bioinform.2
2013 Incorporating fuzzy semantic similarity measure in detecting human protein complexes in PPI network: A multiobjective approach
abstract
Detection of protein complexes within protein-protein interaction networks (PPIN) is a valuable step toward the analysis of biological processes and pathways. Several high-throughput experimental techniques produce large number of PPIs that can be extensively utilized for constructing PPI network of a species. Decomposition of the whole PPI network into smaller and manageable modules is an ongoing challenge. Here we have developed a multi-objective algorithm for detecting human protein complexes by partitioning large human PPI network into clusters which serve as protein complexes. Some graphical properties like density, centrality etc., are utilized for building the objectives. Besides the graphical properties we have also exploited a fuzzy measure based semantic similarity approach to construct similarity based objective. The proposed technique is demonstrated in the human PPI network and the resulting complexes are analyzed in context of Gene Ontology (GO) and pathway enrichment. We have also compared our results with that of some state-of-the-art algorithms in context of different performance metrics. The biological relevance of our predicted complexes are also established here by linking them with 22 key disease classes.
Sumanta Ray, Sanghamitra Bandyopadhyay, Anirban Mukhopadhyay 0001, Ujjwal Maulik
FUZZ-IEEE1