Saurav Mallik

dblp:134/7796 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-4107-6784ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 9 first-author · 14 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Hi-Enhancer: a two-stage framework for prediction and localization of enhancers based on Blending-KAN and Stacking-Auto models
abstract
MOTIVATION: Gene expression plays a crucial role in cell function, and enhancers can regulate gene expression precisely. Therefore, accurate prediction of enhancers is particularly critical. However, existing prediction methods have low accuracy or rely on fixed multiple epigenetic signals, which may not always be available. RESULTS: We propose a two-stage framework that accurately predicts enhancers by flexibly combining multiple epigenetic signals. In the first stage, we designed a Blending-KAN model, which integrates the results of various base classifiers and employs Kolmogorov-Arnold Networks (KAN) as a meta-classifier to predict enhancers based on flexible combinations of multiple epigenetic signals. In the second stage, we developed a Stacking-Auto model, which extracted sequence features using DNABERT-2 and located the enhancers based on the Stacking strategy and AutoGluon framework. The accuracy of the Blending-KAN model reached 99.69 ± 0.11% when five epigenetic signals were used. In cross-cell line prediction, the accuracy was more significant than or equal to 93.72%. With Gaussian noise, it still maintains an accuracy of 98.74 ± 0.03%. In the second stage, the accuracy of the Stacking-Auto model is 80.50%, which is better than the existing 17 methods. The results show that our models can be flexibly used to predict and locate enhancers utilizing a combination of multiple epigenetic signals. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/emanlee/Hi-Enhancer and https://doi.org/10.6084/m9.figshare.29262158.v1.
Rong Fei, Juntao Zou, Xiguo Yuan, Saurav Mallik, Xinhong Hei 0001, Lei Wang 0029
Bioinform.7
2025 DOMSCNet: a deep learning model for the classification of stomach cancer using multi-layer omics data
abstract
The rapid advancement of next-generation sequencing (NGS) technology and the expanding availability of NGS datasets have led to a significant surge in biomedical research. To better understand the molecular processes, underlying cancer and to support its development, diagnosis, prediction, and therapy; NGS data analysis is crucial. However, the NGS multi-layer omics high-dimensional dataset is highly complex. In recent times, some computational methods have been developed for cancer omics data interpretation. However, various existing methods face challenges in accounting for diverse types of cancer omics data and struggle to effectively extract informative features for the integrated identification of core units. To address these challenges, we proposed a hybrid feature selection (HFS) technique to detect optimal features from multi-layer omics datasets. Subsequently, this study proposes a novel hybrid deep recurrent neural network-based model DOMSCNet to classify stomach cancer. The proposed model was made generic for all four multi-layer omics datasets. To observe the robustness of the DOMSCNet model, the proposed model was validated with eight external datasets. Experimental results showed that the SelectKBest-maximum relevancy minimum redundancy-Boruta (SMB), HFS technique outperformed all other HFS techniques. Across four multi-layer omics datasets and validated datasets, the proposed DOMSCNet model outdid existing classifiers along with other proposed classifiers.
Kasmika Borah, Himanish Shekhar Das, Ram Kaji Budhathoki, Khursheed Aurangzeb, Saurav Mallik
Briefings Bioinform.5
2025 Deadline-aware and energy efficient IoT task scheduling using fuzzy logic in fog computing
Rahul Thakur, Geeta Sikka, Urvashi Bansal, Jayant P. Giri, Saurav Mallik
Multim. Tools Appl.5
2024 Predicting stroke occurrences: a stacked machine learning approach with feature selection and data preprocessing
abstract
Stroke prediction remains a critical area of research in healthcare, aiming to enhance early intervention and patient care strategies. This study investigates the efficacy of machine learning techniques, particularly principal component analysis (PCA) and a stacking ensemble method, for predicting stroke occurrences based on demographic, clinical, and lifestyle factors. We systematically varied PCA components and implemented a stacking model comprising random forest, decision tree, and K-nearest neighbors (KNN).Our findings demonstrate that setting PCA components to 16 optimally enhanced predictive accuracy, achieving a remarkable 98.6% accuracy in stroke prediction. Evaluation metrics underscored the robustness of our approach in handling class imbalance and improving model performance, also comparative analyses against traditional machine learning algorithms such as SVM, logistic regression, and Naive Bayes highlighted the superiority of our proposed method.
Pritam Chakraborty, Anjan Bandyopadhyay, Preeti Padma Sahu, Aniket Burman, Saurav Mallik, Mohamed Abbas, Mohammed S. Alqahtani, Ben Othman Soufiene
BMC Bioinform.5
2024 Ribosomal computing: implementation of the computational method
abstract
BACKGROUND: Several computational and mathematical models of protein synthesis have been explored to accomplish the quantitative analysis of protein synthesis components and polysome structure. The effect of gene sequence (coding and non-coding region) in protein synthesis, mutation in gene sequence, and functional model of ribosome needs to be explored to investigate the relationship among protein synthesis components further. Ribosomal computing is implemented by imitating the functional property of protein synthesis. RESULT: In the proposed work, a general framework of ribosomal computing is demonstrated by developing a computational model to present the relationship between biological details of protein synthesis and computing principles. Here, mathematical abstractions are chosen carefully without probing into intricate chemical details of the micro-operations of protein synthesis for ease of understanding. This model demonstrates the cause and effect of ribosome stalling during protein synthesis and the relationship between functional protein and gene sequence. Moreover, it also reveals the computing nature of ribosome molecules and other protein synthesis components. The effect of gene mutation on protein synthesis is also explored in this model. CONCLUSION: The computational model for ribosomal computing is implemented in this work. The proposed model demonstrates the relationship among gene sequences and protein synthesis components. This model also helps to implement a simulation environment (a simulator) for generating protein chains from gene sequences and can spot the problem during protein synthesis. Thus, this simulator can identify a disease that can happen due to a protein synthesis problem and suggest precautions for it.
Pratima Chatterjee, Prasun Ghosal, Sahadeb Shit, Arindam Biswas 0007, Saurav Mallik, Sarah Allabun, Manal Othman, Almubarak Hassan Ali, E. Elshiekh, Ben Othman Soufiene
BMC Bioinform.5
2023 A scalable unsupervised learning of scRNAseq data detects rare cells through integration of structure-preserving embedding, clustering and outlier detection
abstract
Single-cell RNA-seq analysis has become a powerful tool to analyse the transcriptomes of individual cells. In turn, it has fostered the possibility of screening thousands of single cells in parallel. Thus, contrary to the traditional bulk measurements that only paint a macroscopic picture, gene measurements at the cell level aid researchers in studying different tissues and organs at various stages. However, accurate clustering methods for such high-dimensional data remain exiguous and a persistent challenge in this domain. Of late, several methods and techniques have been promulgated to address this issue. In this article, we propose a novel framework for clustering large-scale single-cell data and subsequently identifying the rare-cell sub-populations. To handle such sparse, high-dimensional data, we leverage PaCMAP (Pairwise Controlled Manifold Approximation), a feature extraction algorithm that preserves both the local and the global structures of the data and Gaussian Mixture Model to cluster single-cell data. Subsequently, we exploit Edited Nearest Neighbours sampling and Isolation Forest/One-class Support Vector Machine to identify rare-cell sub-populations. The performance of the proposed method is validated using the publicly available datasets with varying degrees of cell types and rare-cell sub-populations. On several benchmark datasets, the proposed method outperforms the existing state-of-the-art methods. The proposed method successfully identifies cell types that constitute populations ranging from 0.1 to 8% with F1-scores of 0.91 0.09. The source code is available at https://github.com/scrab017/RarPG.
Koushik Mallick, Sikim Chakraborty, Saurav Mallik, Sanghamitra Bandyopadhyay
Briefings Bioinform.3
2023 Detection for melanoma skin cancer through ACCF, BPPF, and CLF techniques with machine learning approach
abstract
Intense sun exposure is a major risk factor for the development of melanoma, an abnormal proliferation of skin cells. Yet, this more prevalent type of skin cancer can also develop in less-exposed areas, such as those that are shaded. Melanoma is the sixth most common type of skin cancer. In recent years, computer-based methods for imaging and analyzing biological systems have made considerable strides. This work investigates the use of advanced machine learning methods, specifically ensemble models with Auto Correlogram Methods, Binary Pyramid Pattern Filter, and Color Layout Filter, to enhance the detection accuracy of Melanoma skin cancer. These results suggest that the Color Layout Filter model of the Attribute Selection Classifier provides the best overall performance. Statistics for ROC, PRC, Kappa, F-Measure, and Matthews Correlation Coefficient were as follows: 90.96% accuracy, 0.91 precision, 0.91 recall, 0.95 ROC, 0.87 PRC, 0.87 Kappa, 0.91 F-Measure, and 0.82 Matthews Correlation Coefficient. In addition, its margins of error are the smallest. The research found that the Attribute Selection Classifier performed well when used in conjunction with the Color Layout Filter to improve image quality.
G. Ayyappan, Prabhu Jayagopal, Sandeep Kumar Mathivanan 0001, Saurav Mallik, Amal Al-Rasheed, Mohammed S. Alqahtani, Ben Othman Soufiene
BMC Bioinform.5
2023 A novel and innovative cancer classification framework through a consecutive utilization of hybrid feature selection
abstract
Cancer prediction in the early stage is a topic of major interest in medicine since it allows accurate and efficient actions for successful medical treatments of cancer. Mostly cancer datasets contain various gene expression levels as features with less samples, so firstly there is a need to eliminate similar features to permit faster convergence rate of classification algorithms. These features (genes) enable us to identify cancer disease, choose the best prescription to prevent cancer and discover deviations amid different techniques. To resolve this problem, we proposed a hybrid novel technique CSSMO-based gene selection for cancer classification. First, we made alteration of the fitness of spider monkey optimization (SMO) with cuckoo search algorithm (CSA) algorithm viz., CSSMO for feature selection, which helps to combine the benefit of both metaheuristic algorithms to discover a subset of genes which helps to predict a cancer disease in early stage. Further, to enhance the accuracy of the CSSMO algorithm, we choose a cleaning process, minimum redundancy maximum relevance (mRMR) to lessen the gene expression of cancer datasets. Next, these subsets of genes are classified using deep learning (DL) to identify different groups or classes related to a particular cancer disease. Eight different benchmark microarray gene expression datasets of cancer have been utilized to analyze the performance of the proposed approach with different evaluation matrix such as recall, precision, F1-score, and confusion matrix. The proposed gene selection method with DL achieves much better classification accuracy than other existing DL and machine learning classification models with all large gene expression dataset of cancer.
Rajul Mahto, Saboor Uddin Ahmed, Rizwan Ur Rahman, Rabia Musheer Aziz, Saurav Mallik
BMC Bioinform.6
2023 Multimodal hybrid convolutional neural network based brain tumor grade classification
abstract
An abnormal growth or fatty mass of cells in the brain is called a tumor. They can be either healthy (normal) or become cancerous, depending on the structure of their cells. This can result in increased pressure within the cranium, potentially causing damage to the brain or even death. As a result, diagnostic procedures such as computed tomography, magnetic resonance imaging, and positron emission tomography, as well as blood and urine tests, are used to identify brain tumors. However, these methods can be labor-intensive and sometimes yield inaccurate results. Instead of these time-consuming methods, deep learning models are employed because they are less time-consuming, require less expensive equipment, produce more accurate results, and are easy to set up. In this study, we propose a method based on transfer learning, utilizing the pre-trained VGG-19 model. This approach has been enhanced by applying a customized convolutional neural network framework and combining it with pre-processing methods, including normalization and data augmentation. For training and testing, our proposed model used 80% and 20% of the images from the dataset, respectively. Our proposed method achieved remarkable success, with an accuracy rate of 99.43%, a sensitivity of 98.73%, and a specificity of 97.21%. The dataset, sourced from Kaggle for training purposes, consists of 407 images, including 257 depicting brain tumors and 150 without tumors. These models could be utilized to develop clinically useful solutions for identifying brain tumors in CT images based on these outcomes.
A. Rohini, Carol Praveen, Sandeep Kumar Mathivanan 0001, V. Muthukumaran, Saurav Mallik, Mohammed S. Alqahtani, Amal Al-Rasheed, Ben Othman Soufiene
BMC Bioinform.5
2023 AngClust: Angle Feature-Based Clustering for Short Time Series Gene Expression Profiles
abstract
When clustering gene expression, it is expected that correlation coefficients of genes in the same clusters are high, and that gene ontology (GO) enrichment analysis of most clusters will be significant. However, existing short-term gene expression clustering algorithms have limitations. To address this problem, we proposed a novel clustering process based on angular features for short-term gene expression. Our method (named AngClust) uses angular features to indicate the change of trend in gene expression levels at two neighboring time points. The changes of angles at multiple time points reflects the change of trend of the overall expression levels. Such changes are used to measure whether the expression trends of different genes are similar. To obtain functionally significant clusters from the clustering results, we evaluated numbers of genes in clusters, average correlation coefficient, fluctuation, and their correlation with GO term enrichment. The efficacy of AngClust outperform two other measures, Euclidean distance (ED) and dynamic time warping of correlation (DTW), on a dataset of yeast gene expression. The ratios of GO and pathway term-enriched of clusters of AngClust is higher than or equal to that of STEM and TMixClust on human, mouse, and yeast time series of gene expression.
Junhuai Li, Saurav Mallik, Rong Fei, Hongfang Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 Detection of Chronic Kidney Disease Using Neuro-Fuzzy Rule-based Classifier
abstract
Chronic kidney disease (CKD) is as severe as cancer in today s world. It may even lead to the permanent failure of kidney. The initial detection of this disease is needed for timely cure. Our work presents a classifier (named ANFIS) in accordance with the notion of neuro-fuzzy in order to detect the existence of chronic kidney disease. We use blood test results of several patients for our research study. We compare our proposed classifier with some conventional classifiers such as Multi-layer Perceptron, Support Vector Machine, Logistic Regression and Decision Tree. Experimental results indicates that our proposed neuro-fuzzy rule-based classifier performs better than the other classifiers used here. ANFIS has given 3% to 4% better accuracy compared to the other classifiers.
Supantha Das, Arnab Hazra, Soumen Kumar Pati, Soumadip Ghosh, Saurav Mallik, Suharta Banerjee, Ayan Mukherji, Zhongming Zhao
BIBM5
2022 Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous Processor
abstract
RF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design.
Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue
ISCAS30
2022 Comparison of five supervised feature selection algorithms leading to top features and gene signatures from multi-omics data in cancer
abstract
BACKGROUND: As many complex omics data have been generated during the last two decades, dimensionality reduction problem has been a challenging issue in better mining such data. The omics data typically consists of many features. Accordingly, many feature selection algorithms have been developed. The performance of those feature selection methods often varies by specific data, making the discovery and interpretation of results challenging. METHODS AND RESULTS: In this study, we performed a comprehensive comparative study of five widely used supervised feature selection methods (mRMR, INMIFS, DFS, SVM-RFE-CBR and VWMRmR) for multi-omics datasets. Specifically, we used five representative datasets: gene expression (Exp), exon expression (ExpExon), DNA methylation (hMethyl27), copy number variation (Gistic2), and pathway activity dataset (Paradigm IPLs) from a multi-omics study of acute myeloid leukemia (LAML) from The Cancer Genome Atlas (TCGA). The different feature subsets selected by the aforesaid five different feature selection algorithms are assessed using three evaluation criteria: (1) classification accuracy (Acc), (2) representation entropy (RE) and (3) redundancy rate (RR). Four different classifiers, viz., C4.5, NaiveBayes, KNN, and AdaBoost, were used to measure the classification accuary (Acc) for each selected feature subset. The VWMRmR algorithm obtains the best Acc for three datasets (ExpExon, hMethyl27 and Paradigm IPLs). The VWMRmR algorithm offers the best RR (obtained using normalized mutual information) for three datasets (Exp, Gistic2 and Paradigm IPLs), while it gives the best RR (obtained using Pearson correlation coefficient) for two datasets (Gistic2 and Paradigm IPLs). It also obtains the best RE for three datasets (Exp, Gistic2 and Paradigm IPLs). Overall, the VWMRmR algorithm yields best performance for all three evaluation criteria for majority of the datasets. In addition, we identified signature genes using supervised learning collected from the overlapped top feature set among five feature selection methods. We obtained a 7-gene signature (ZMIZ1, ENG, FGFR1, PAWR, KRT17, MPO and LAT2) for EXP, a 9-gene signature for ExpExon, a 7-gene signature for hMethyl27, one single-gene signature (PIK3CG) for Gistic2 and a 3-gene signature for Paradigm IPLs. CONCLUSION: We performed a comprehensive comparison of the performance evaluation of five well-known feature selection methods for mining features from various high-dimensional datasets. We identified signature genes using supervised learning for the specific omic data for the disease. The study will help incorporate higher order dependencies among features.
Tapas Bhadra, Saurav Mallik, Neaj Hasan, Zhongming Zhao
BMC Bioinform.2
2022 Unsupervised Feature Selection Using an Integrated Strategy of Hierarchical Clustering With Singular Value Decomposition: An Integrative Biomarker Discovery Method With Application to Acute Myeloid Leukemia
abstract
In this article, we propose a novel unsupervised feature selection method by combining hierarchical feature clustering with singular value decomposition (SVD). The proposed algorithm first generates several feature clusters by adopting the hierarchical clustering on the feature space and then applies SVD to each of these feature clusters to find out the feature that contributes most to the SVD-entropy. The proposed feature selection method selects an optimal feature subset that not only minimizes the mutual dependency among the selected features but also maximizes the mutual dependency of the selected features against their nearest neighbor non-selected features to some extent. Each of the selected features also contributes the maximum SVD-entropy among all features of the same feature cluster. The experimental results demonstrate that the proposed algorithm performs well against several state-of-the-art methods of feature selection in terms of various evaluation criteria such as classification accuracy, redundancy rate, and representation entropy. The superiority of the proposed algorithm is demonstrated through analysis of Acute Myeloid Leukemia (AML) multi-omics data that consist of five datasets: gene expression, exon expression, methylation, microRNA, and pathway activity dataset (paradigm IPLs) from The Cancer Genome Atlas (TCGA). Our analysis pinpoints a candidate gene-marker, EREG for AML with an integrative omics evidence. EREG is targeted by two top ranked microRNAs, hsa-miR-1286 and hsa-miR-1976, here in the datasets. The method and results will be useful for biomarker discovery in the era of in precision medicine.
Tapas Bhadra, Saurav Mallik, Amir Sohel, Zhongming Zhao
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 A Novel Graph Topology-Based GO-Similarity Measure for Signature Detection From Multi-Omics Data and its Application to Other Problems
abstract
Large scale multi-omics data analysis and signature prediction have been a topic of interest in the last two decades. While various traditional clustering/correlation-based methods have been proposed, but the overall prediction is not always satisfactory. To solve these challenges, in this article, we propose a new approach by leveraging the Gene Ontology (GO)similarity combined with multiomics data. In this article, a new GO similarity measure, ModSchlicker, is proposed and the effectiveness of the proposed measure along with other standardized measures are reviewed while using various graph topology-based Information Content (IC)values of GO-term. The proposed measure is deployed to PPI prediction. Furthermore, by involving GO similarity, we propose a new framework for stronger disease-based gene signature detection from the multi-omics data. For the first objective, we predict interaction from various benchmark PPI datasets of Yeast and Human species. For the latter, the gene expression and methylation profiles are used to identify Differentially Expressed and Methylated (DEM)genes. Thereafter, the GO similarity score along with a statistical method are used to determine the potential gene signature. Interestingly, the proposed method produces a better performance ( 0.9 avg. accuracy and 0.95 AUC)as compared to the other existing related methods during the classification of the participating features (genes)of the signature. Moreover, the proposed method is highly useful in other prediction/classification problems for any kind of large scale omics data.
Koushik Mallick, Saurav Mallik, Sanghamitra Bandyopadhyay, Sikim Chakraborty
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Negatively-Associated Maximal Frequent Geneset Mining on DNA Methylation Profile
abstract
Association rule mining has been an important approach for feature and biomarker discovery in various omics data. One main challenge is that it generates a large number of itemsets. The effect of this shortcoming increases substantially in the case of negative association rule mining that is useful for detecting significant relationships between genes (items/features) in the form of either presence or absence in disease characterization. In this article, we propose a new algorithm, NegaMax (negatively-associated maximal frequent itemsets) for negative association itemset mining. Our method follows depth-first search rather than breadth-first search used in the other methods. It identifies a much fewer number of non-redundant itemsets than that by the existing methods. Thus, it saves elapsing time for itemset generation which potentially remove false positive intermediate results. We demonstrated NegaMax in a real-world DNA methylation dataset. The proposed method is highly beneficial from a medical perspective.
Saurav Mallik, Souvik Rakshit, Ujjwal Maulik, Zhongming Zhao
BIBM1
2020 Critical microRNAs and regulatory motifs in cleft palate identified by a conserved miRNA-TF-gene network approach in humans and mice
abstract
Cleft palate (CP) is the second most common congenital birth defect. The etiology of CP is complicated, with involvement of various genetic and environmental factors. To investigate the gene regulatory mechanisms, we designed a powerful regulatory analytical approach to identify the conserved regulatory networks in humans and mice, from which we identified critical microRNAs (miRNAs), target genes and regulatory motifs (miRNA-TF-gene) related to CP. Using our manually curated genes and miRNAs with evidence in CP in humans and mice, we constructed miRNA and transcription factor (TF) co-regulation networks for both humans and mice. A consensus regulatory loop (miR17/miR20a-FOXE1-PDGFRA) and eight miRNAs (miR-140, miR-17, miR-18a, miR-19a, miR-19b, miR-20a, miR-451a and miR-92a) were discovered in both humans and mice. The role of miR-140, which had the strongest association with CP, was investigated in both human and mouse palate cells. The overexpression of miR-140-5p, but not miR-140-3p, significantly inhibited cell proliferation. We further examined whether miR-140 overexpression could suppress the expression of its predicted target genes (BMP2, FGF9, PAX9 and PDGFRA). Our results indicated that miR-140-5p overexpression suppressed the expression of BMP2 and FGF9 in cultured human palate cells and Fgf9 and Pdgfra in cultured mouse palate cells. In summary, our conserved miRNA-TF-gene regulatory network approach is effective in detecting consensus miRNAs, motifs, and regulatory mechanisms in human and mouse CP.
Peilin Jia, Saurav Mallik, Rong Fei, Hiroki Yoshioka, Akiko Suzuki, Junichi Iwata, Zhongming Zhao
Briefings Bioinform.3
2020 Graph- and rule-based learning algorithms: a comprehensive review of their applications for cancer type classification and prognosis using genomic data
abstract
Cancer is well recognized as a complex disease with dysregulated molecular networks or modules. Graph- and rule-based analytics have been applied extensively for cancer classification as well as prognosis using large genomic and other data over the past decade. This article provides a comprehensive review of various graph- and rule-based machine learning algorithms that have been applied to numerous genomics data to determine the cancer-specific gene modules, identify gene signature-based classifiers and carry out other related objectives of potential therapeutic value. This review focuses mainly on the methodological design and features of these algorithms to facilitate the application of these graph- and rule-based analytical approaches for cancer classification and prognosis. Based on the type of data integration, we divided all the algorithms into three categories: model-based integration, pre-processing integration and post-processing integration. Each category is further divided into four sub-categories (supervised, unsupervised, semi-supervised and survival-driven learning analyses) based on learning style. Therefore, a total of 11 categories of methods are summarized with their inputs, objectives and description, advantages and potential limitations. Next, we briefly demonstrate well-known and most recently developed algorithms for each sub-category along with salient information, such as data profiles, statistical or feature selection methods and outputs. Finally, we summarize the appropriate use and efficiency of all categories of graph- and rule mining-based learning methods when input data and specific objective are given. This review aims to help readers to select and use the appropriate algorithms for cancer classification and prognosis study.
Saurav Mallik, Zhongming Zhao
Briefings Bioinform.1
2020 In silico ranking of phenolics for therapeutic effectiveness on cancer stem cells
abstract
BACKGROUND: Cancer stem cells (CSCs) have features such as the ability to self-renew, differentiate into defined progenies and initiate the tumor growth. Treatments of cancer include drugs, chemotherapy and radiotherapy or a combination. However, treatment of cancer by various therapeutic strategies often fail. One possible reason is that the nature of CSCs, which has stem-like properties, make it more dynamic and complex and may cause the therapeutic resistance. Another limitation is the side effects associated with the treatment of chemotherapy or radiotherapy. To explore better or alternative treatment options the current study aims to investigate the natural drug-like molecules that can be used as CSC-targeted therapy. Among various natural products, anticancer potential of phenolics is well established. We collected the 21 phytochemicals from phenolic group and their interacting CSC genes from the publicly available databases. Then a bipartite graph is constructed from the collected CSC genes along with their interacting phytochemicals from phenolic group as other. The bipartite graph is then transformed into weighted bipartite graph by considering the interaction strength between the phenolics and the CSC genes. The CSC genes are also weighted by two scores, namely, DSI (Disease Specificity Index) and DPI (Disease Pleiotropy Index). For each gene, its DSI score reflects the specific relationship with the disease and DPI score reflects the association with multiple diseases. Finally, a ranking technique is developed based on PageRank (PR) algorithm for ranking the phenolics. RESULTS: We collected 21 phytochemicals from phenolic group and 1118 CSC genes. The top ranked phenolics were evaluated by their molecular and pharmacokinetics properties and disease association networks. We selected top five ranked phenolics (Resveratrol, Curcumin, Quercetin, Epigallocatechin Gallate, and Genistein) for further examination of their oral bioavailability through molecular properties, drug likeness through pharmacokinetic properties, and associated network with CSC genes. CONCLUSION: Our PR ranking based approach is useful to rank the phenolics that are associated with CSC genes. Our results suggested some phenolics are potential molecules for CSC-related cancer treatment.
Monalisa Mandal, Sanjeeb Kumar Sahoo, Priyadarsan Patra, Saurav Mallik, Zhongming Zhao
BMC Bioinform.4
2020 Determining the interaction status and evolutionary fate of duplicated homomeric proteins
abstract
Oligomeric proteins are central to life. Duplication and divergence of their genes is a key evolutionary driver, also because duplications can yield very different outcomes. Given a homomeric ancestor, duplication can yield two paralogs that form two distinct homomeric complexes, or a heteromeric complex comprising both paralogs. Alternatively, one paralog remains a homomer while the other acquires a new partner. However, so far, conflicting trends have been noted with respect to which fate dominates, primarily because different methods and criteria are being used to assign the interaction status of paralogs. Here, we systematically analyzed all Saccharomyces cerevisiae and Escherichia coli oligomeric complexes that include paralogous proteins. We found that the proportions of homo-hetero duplication fates strongly depend on a variety of factors, yet that nonetheless, rigorous filtering gives a consistent picture. In E. coli about 50%, of the paralogous pairs appear to have retained the ancestral homomeric interaction, whereas in S. cerevisiae only ~10% retained a homomeric state. This difference was also observed when unique complexes were counted instead of paralogous gene pairs. We further show that this difference is accounted for by multiple cases of heteromeric yeast complexes that share common ancestry with homomeric bacterial complexes. Our analysis settles contradicting trends and conflicting previous analyses, and provides a systematic and rigorous pipeline for delineating the fate of duplicated oligomers in any organism for which protein-protein interaction data are available.
Saurav Mallik, Dan S. Tawfik
PLoS Comput. Biol.1
2020 WeCoMXP: Weighted Connectivity Measure Integrating Co-Methylation, Co-Expression and Protein-Protein Interactions for Gene-Module Detection
abstract
The identification of modules (groups of several tightly interconnected genes) in gene interaction network is an essential task for better understanding of the architecture of the whole network. In this article, we develop a novel weighted connectivity measure integrating co-methylation, co-expression, and protein-protein interactions (called WeCoMXP) to detect gene-modules for multi-omics dataset. The proposed measure goes beyond the fundamental degree centrality measure through considering some formulation of higher-order connections. Thereafter, we apply the average linkage clustering method using the corresponding dissimilarity (distance) values of WeCoMXP scores, and utilize a dynamic tree cut method for identifying some gene-modules. We validate the modules through literature search, KEGG pathway, and gene-ontology analyses on the genes representing the modules. Furthermore, the top 10 TFs/miRNAs that are connected with the maximum number of gene-modules and that regulate/target the maximum number of genes from these connected gene-modules, are identified. Moreover, our proposed method provides a better performance than the existing methods in terms of several cluster-validity indices in maximum times.
Saurav Mallik, Sanghamitra Bandyopadhyay
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 A Multi-classifier Model to Identify Mitochondrial Respiratory Gene Signatures in Human Cancer
abstract
Whether alteration of mitochondrial gene expression can serve as an effective molecular signature for cancer classification currently remains controversy. To tackle this challenge, here we present a multi-classifier model to identify mitochondrial aberrant gene signatures and then assess the effectiveness on cancer classification. Specifically, we first applied a supervised learning model, Empirical Bayes statistics, to detect differentially expressed genes from a liver cancer mitochondrial gene expression dataset (GEO accession: GSE64505). Next, we applied two well-known classifiers, Prediction Analysis of Microarrays (PAM) and Random Forest (RF), with several folds of cross-validation, to find molecular signatures that can classify liver cancer samples from normal samples. We obtained a mitochondrial molecular signature comprising 25 genes. Classification accuracy was 87.5% by PAM classifier (5 fold cross-validation with 30 repeats), while it was 87.5% and 75.0% by Random Forest with 5- and 2-fold cross-validations (30 repeats), respectively. We further performed literature mining and Gene Set Enrichment Analysis (GSEA) to evaluate the biological significance and novelty of genes in this gene signature.
Saurav Mallik, Soumita Seth, Tapas Bhadra, Namrata Tomar, Zhongming Zhao
BIBM1
2019 An evaluation of supervised methods for identifying differentially methylated regions in Illumina methylation arrays
abstract
Epigenome-wide association studies (EWASs) have become increasingly popular for studying DNA methylation (DNAm) variations in complex diseases. The Illumina methylation arrays provide an economical, high-throughput and comprehensive platform for measuring methylation status in EWASs. A number of software tools have been developed for identifying disease-associated differentially methylated regions (DMRs) in the epigenome. However, in practice, we found these tools typically had multiple parameter settings that needed to be specified and the performance of the software tools under different parameters was often unclear. To help users better understand and choose optimal parameter settings when using DNAm analysis tools, we conducted a comprehensive evaluation of 4 popular DMR analysis tools under 60 different parameter settings. In addition to evaluating power, precision, area under precision-recall curve, Matthews correlation coefficient, F1 score and type I error rate, we also compared several additional characteristics of the analysis results, including the size of the DMRs, overlap between the methods and execution time. The results showed that none of the software tools performed best under their default parameter settings, and power varied widely when parameters were changed. Overall, the precision of these software tools were good. In contrast, all methods lacked power when effect size was consistent but small. Across all simulation scenarios, comb-p consistently had the best sensitivity as well as good control of false-positive rate.
Saurav Mallik, Gabriel J. Odom, Lissette Gomez, Lily Wang 0001
Briefings Bioinform.1
2019 Identification of Multiview Gene Modules Using Mutual Information-Based Hypograph Mining
abstract
Detection of gene-modules is one of the fundamental tasks for the integral analysis of network architecture. In this paper, we propose a novel algorithm using an integrated approach comprising statistical method and normalized mutual information-based hypograph mining for discovering the multiview co-similarity gene modules contained in multiview datasets. For this purpose, we first identify the statistically significant genes corresponding to each data profile and subsequently obtain the union set consisting of all these statistically significant genes. For each data profile, we then propose a new similarity score called as integrated normalized mutual information to obtain the similarity scores across all possible pairs of genes belonging to the union set by employing the results of gene clustering obtained through applying normalized mutual information-based graph clustering on the corresponding data profile. Moreover, we propose a new information theoretic measure called as multiview normalized mutual information to integrate all the similarity scores of a given gene-pair obtained across all the data profiles. For the experiment, we utilize one of the recently used multiview dataset named TCGA acute myeloid leukemia dataset comprising five different categories of data profiles. Furthermore, the co-similarity strengths of all the multiview gene modules obtained using the proposed method (PM) are reported. Finally, we provide a comparative study between the proposed and other existing methods for demonstrating the superiority of the PM over others. Code is available in http://ieeexplore.ieee.org.
Tapas Bhadra, Saurav Mallik, Sanghamitra Bandyopadhyay
IEEE Trans. Syst. Man Cybern. Syst.2
2018 Integrating Multiple Data Sources for Combinatorial Marker Discovery: A Study in Tumorigenesis
Sanghamitra Bandyopadhyay, Saurav Mallik
IEEE ACM Trans. Comput. Biol. Bioinform.2
2017 TrapRM: Transcriptomic and proteomic rule mining using weighted shortest distance based multiple minimum supports for multi-omics dataset
abstract
Association rule mining is an important machine learning tool for unveiling critical biological relations between genes from omics data. Previous approaches typically are designed for one single genomic dataset, and most of them use a single minimum support threshold globally. To overcome the above two general limitations, in this work, we present a novel Transcriptomic and Proteomic Rule Mining (TrapRM) method using Weighted Shortest Distance based Multiple Minimum Supports for Multi-Omics Dataset that integrates gene expression, methylation and protein-protein interaction data. To do so, we initially introduce three new thresholds: Weighted Shortest Distance based Multiple Minimum Supports (WSDMS), Weighted Shortest Distance based Multiple Minimum Confidences (WSDMC), and Weighted Shortest Distance based Multiple Minimum Lifts (WSDML). Our algorithm is superior to the related existing algorithms since it generates substantially fewer number of rules and smaller average weighted shortest distance value than the existing methods. Finally, our TrapRM algorithm is useful for extracting the rules that are critical for translational and clinical applications when being applied to drug or disease related multi-omics data.
Saurav Mallik, Zhongming Zhao
BIBM1
2015 MiRNA-TF-gene network analysis through ranking of biomolecules for multi-informative uterine leiomyoma dataset
Saurav Mallik, Ujjwal Maulik
J. Biomed. Informatics1
2014 A Survey and Comparative Study of Statistical Tests for Identifying Differential Expressionfrom Microarray Data
abstract
DNA microarray is a powerful technology that can simultaneously determine the levels of thousands of transcripts (generated, for example, from genes/miRNAs) across different experimental conditions or tissue samples. The motto of differential expression analysis is to identify the transcripts whose expressions change significantly across different types of samples or experimental conditions. A number of statistical testing methods are available for this purpose. In this paper, we provide a comprehensive survey on different parametric and non-parametric testing methodologies for identifying differential expression from microarray data sets. The performances of the different testing methods have been compared based on some real-life miRNA and mRNA expression data sets. For validating the resulting differentially expressed miRNAs, the outcomes of each test are checked with the information available for miRNA in the standard miRNA database PhenomiR 2.0. Subsequently, we have prepared different simulated data sets of different sample sizes (from 10 to 100 per group/population) and thereafter the power of each test have been calculated individually. The comparative simulated study might lead to formulate robust and comprehensive judgements about the performance of each test in the basis of assumption of data distribution. Finally, a list of advantages and limitations of the different statistical tests has been provided, along with indications of some areas where further studies are required.
Sanghamitra Bandyopadhyay, Saurav Mallik, Anirban Mukhopadhyay 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2013 Integrated analysis of gene expression and genome-wide DNA methylation for tumor prediction: An association rule mining-based approach
abstract
Statistical analysis and association rule mining are two most efficient techniques, where the first one is used to identify differentially expressed/methylated genes across different types of samples or experimental conditions and the second one is used to determine expression/methylation relationships among them. In this article, we have performed an integrated analysis of statistical methods and association rule mining on mRNA expression and DNA methylation datasets for the prediction of Uterine Leiomyoma. Moreover, we have proposed a novel rule-base classifier. Depending on 16 different rule-interestingness measures, we have applied a Genetic Algorithm based rank aggregation technique on the association rules which are generated from the training data by Apriori association rule mining algorithm. After determining the ranks of the rules, we have conducted a majority voting technique on each test point to determine its class-label (i.e. tumor or normal class-label) through weighted-sum method. We have run this classifier on the combined dataset using k-fold cross-validation and also performed a comparative performance analysis with other popular rule-base classifiers. Finally, we have predicted the status of some important genes (through frequency analysis in association rules for tumor and normal class-labels individually) that have a major role for tumor formation in Uterine Leiomyoma.
Saurav Mallik, Anirban Mukhopadhyay 0001, Ujjwal Maulik, Sanghamitra Bandyopadhyay
CIBCB1