M. Mondal Ananda

dblp:53/9077 · also Ananda Mohan Mondal, Ananda Mondal · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0002-4005-9942ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2024 REDONE-PD: Reflections of Dopamine-Related Gene Mutations on Neurocognitive Functions in Healthy Controls and Parkinson's Disease
abstract
Parkinson’s Disease (PD) is a neurodegenerative disorder characterized by both motor and non-motor symptoms, including significant changes in neurocognitive functions (NFs). Dopamine synthesis, a critical process in PD, is heavily impacted by genetic factors, contributing to motor dysfunction and neurocognitive impairment. While some genetic mutations have been linked to PD-related neurocognitive impairments, the specific impact of these mutations on NFs remains unclear. This study explores the relationship between mutations in dopamine synthesis-related genes and NFs in PD patients and healthy controls (HC). Using a gene sequencing dataset (INDELs and SNPs) from the PPMI repository that includes 171 dopamine synthesis-related genes, we applied t-tests to identify 34 significantly mutated genes. We considered subjective (self-reported) responses from the MDS-UPDRS for seven NFs (i.e., motor, autonomic function, behavior/psychological, executive function, sensory, sleep, and speech). To investigate the link between significantly mutated genes and neurocognitive performance, participants were grouped into healthy and PD groups, and then each group was split into mutated and non-mutated samples. For healthy controls, presence of mutation in CAMK2D, and NOSTRIN and absence of mutation in DDX17 found related to NF worsening. In case of PD samples, the presence of mutations in genes like B3GALT5, CHRNA2, SNX27, TRIP4, and UBQLN4 and absence of mutation in genes like ABCA7, INTS4, MYLK3, PRKG1, and PRKN were linked to worsening neurocognitive outcomes. Overall, these findings provide a deeper understanding of the genetic influences on neurocognitive functions in PD and highlight potential targets for future research and therapeutic strategies.
Md Mezbahul Islam, John Michael Templeton, Christian Poellabauer, M. Mondal Ananda
BIBM4
2024 Interpreting Lung Cancer Health Disparity between African American Males and European American Males
abstract
Lung cancer remains a predominant cause of cancer-related deaths, with notable disparities in incidence and outcomes across racial and gender groups. This study addresses these disparities by developing a computational framework leveraging explainable artificial intelligence (XAI) to identify both patient- and cohort-specific biomarker genes in lung cancer. Specifically, we focus on two lung cancer subtypes, Lung Adenocarcinoma (LUAD) and Lung Squamous Cell Carcinoma (LUSC), examining distinct racial and sex-specific cohorts: African American males (AAMs) and European American males (EAMs). This study innovatively structures classification tasks based on disease conditions rather than racial labels to avoid race-specific imbalance. We constructed four classification tasks- one three-class problem (LUAD-LUSC-HEALTHY) and three two-class problems (LUAD-LUSC, LUAD-HEALTHY, LUSC-HEALTHY)- to interpret the disease behavior of the patients in terms of genes and pathways. This methodology allows a LUAD or LUSC patient to be analyzed via multiple classifications, yielding robust disparity information for every patient. This preliminary work reports the disparity information for LUAD only. Utilizing Transcriptome data from The Cancer Genome Atlas (TCGA) and Genotype-Tissue Expression (GTEx) projects, we processed samples for LUAD, LUSC, and HEALTHY cohorts. We applied machine learning models, including convolutional neural network (CNN), logistic regression (LR), naïve Bayesian classifier (NB), support vector machine (SVM), random forest (RF), and extreme gradient boosting (XGBoost) for the classification. The SHapley Additive exPlanation (SHAP)-based interpretation of the best performing classification model uncovered cohort-specific genes and pathways related to health disparities between LUAD-AAM and LUAD-EAM cohorts.
Masrur Sobhan, Md Mezbahul Islam, M. Mondal Ananda
BIBM3
2024 Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have become a building block in graph data processing, with wide applications in critical domains. The growing needs to deploy GNNs in high-stakes applications necessitate explainability for users in the decision-making processes. A popular paradigm for the explainability of GNNs is to identify explainable subgraphs by comparing their labels with the ones of original graphs. This task is challenging due to the substantial distributional shift from the original graphs in the training set to the set of explainable subgraphs, which prevents accurate prediction of labels with the subgraphs. To address it, in this paper, we propose a novel method that generates proxy graphs for explainable subgraphs that are in the distribution of training data. We introduce a parametric method that employs graph generators to produce proxy graphs. A new training objective based on information theory is designed to ensure that proxy graphs not only adhere to the distribution of training data but also preserve explanatory factors. Such generated proxy graphs can be reliably used to approximate the predictions of the labels of explainable subgraphs. Empirical evaluations across various datasets demonstrate our method achieves more accurate explanations for GNNs.
Zhuomin Chen, Jiaxing Zhang 0002, Jingchao Ni, Yuchen Bian, Md Mezbahul Islam, M. Mondal Ananda, Hua Wei 0001
ICML7
2023 Evaluating SHAP's Robustness in Precision Medicine: Effect of Filtering and Normalization
abstract
Local interpretation of explainable AI, SHAP (SHapley Additive exPlanations), in disease classification problems offers significant feature scores for each sample, potentially identifying precision medicine targets. Tailoring treatments based on individual genetic and molecular targets can enhance therapeutic outcomes while minimizing side effects. However, the suitability of SHAP's local interpretation at the patient level remains uncertain. It generates different sets of patient-specific genes in various runs, even with consistent overall accuracies. This uncertainty challenges the reliability of SHAP’s local interpretations for precision medicine applications. Not only that, different filtering criteria and normalization techniques may influence the contribution scores of patient-specific features. To validate our hypothesis, SHAP was applied to machine learning algorithms from different genres to identify patient-specific feature contributions from the breast cancer subtype classification problem. The program underwent multiple runs to assess the robustness of SHAP.Our study demonstrates that shallow machine learning algorithms, like Logistic Regression, consistently provided stable and reliable results across multiple runs. In contrast, complex machine learning models like XGBoost and MLP exhibited inconsistencies across different runs. Moreover, we found that data normalization techniques, particularly z-score and min-max normalization, had a minimal effect on the performance of XGBoost models. Our study also shows that the accuracy scores of complex machine learning models remained relatively constant across different runs but produced different sets of patient-specific features. In conclusion, our findings underscore the importance of selecting appropriate filtering and normalization techniques, given the variability in SHAP results across different runs. Our study indicates that combining SHAP with shallow machine learning algorithms yields more stable and dependable results compared to complex machine learning approaches.
Masrur Sobhan, M. Mondal Ananda
BIBM2
2022 Explainable Machine Learning to Identify Patient-specific Biomarkers for Lung Cancer
abstract
Background: Lung cancer is the leading cause of death compared to other cancers in the USA. The overall survival rate of lung cancer is not satisfactory even though there are cutting-edge treatment methods for cancers. Genomic profiling and biomarker gene identification of lung cancer patients may play a role in the therapeutics of lung cancer patients. The biomarker genes identified by most of the existing methods (statistical and machine learning based) belong to the whole cohort or population. That is why different people with the same disease get the same kind of treatment, but results in different outcomes in terms of success and side effects. So, the identification of biomarker genes for individual patients is very crucial for finding efficacious therapeutics leading to precision medicine. Methods: In this study, we propose a pipeline to identify lung cancer class-specific and patient-specific key genes which may help formulate effective therapies for lung cancer patients. We have used expression profiles of two types of lung cancers, lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC), and Healthy lung tissues to identify LUADand LUSC-specific (class-specific) and individual patient-specific key genes using an explainable machine learning approach, SHaphley Additive ExPlanations (SHAP). This approach provides scores for each of the genes for individual patients which tells us the attribution of each feature (gene) for each sample (patient). Result: In this study, we applied two variations of SHAP- tree explainer and gradient explainer for which tree-based classifier, XGBoost, and deep learning-based classifier, convolutional neural network (CNN) were used as classification algorithms, respectively. Our results showed that the proposed approach successfully identified class-specific (LUAD, LUSC, and Healthy) and patient-specific key genes based on the SHAP scores. Conclusion: This study demonstrated a pipeline to identify cohort-based and patient-specific biomarker genes by incorporating an explainable machine learning technique, SHAP. The patient-specific genes identified using SHAP scores may provide biological and clinical insights into the patient’s diagnosis.
Masrur Sobhan, M. Mondal Ananda
BIBM2
2022 An Autoencoder Based Bioinformatics Framework for Predicting Prognosis of Breast Cancer Patients
abstract
It is crucial to find prognostic biomarkers that can predict the cancer prognosis and estimate risk, as they can be used in clinical settings to treat patients. Probing the biomarkers themselves will reveal important insights into the cancer dynamics and molecular pathways underlying pathological behavior. To achieve that goal, this work proposes a bioinformatics framework, taking advantage of the deep learning-based feature selection method Concrete Autoencoder (CAE) to identify key genes and to build a prognostic score model that can assess the risk of cancer patients. 48 gene-pairs were identified to form a prognostic signature model that can significantly differentiate between high-risk and low-risk patients with breast cancer. This prognostic signature was comprised of 42 genes enriched in cancer-related pathways and molecular functions. The proposed framework and the prognostic model can be used as clinical tools to assess the risk levels of breast cancer patients. The identified genes can be studied further for potential targets for cancer therapy.
Raihanul Bari Tanvir, Masrur Sobhan, M. Mondal Ananda
BIBM3
2020 Pan-cancer Feature Selection and Classification Reveals Important Long Non-coding RNAs
abstract
Long noncoding RNA plays important role in changing the expression profiles of various target genes that leads to cancer development. So, identifying key lncRNAs related to the origin of different types of cancers might help in developing cancer therapy. To discover the critical lncRNAs that can identify the origin of different cancers, we proposed to use the state-of-the-art deep learning algorithm Concreate Autoencoder (CAE). The motivation behind using the CAE was that it takes advantage of both AE (which can achieve the highest classification accuracy) and concrete relaxation-based feature selection (which is capable of selecting actual features instead of latent features). To compare the performance of CAE, three frequently used embedded feature selection techniques including Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest (RF), and Support Vector Machine with Recursive Feature Elimination (SVM-RFE) were used. To obtain a stable set of lncRNAs capable of identifying the origin of 33 different cancers, a lncRNA that was isolated by at least two of the four techniques (CAE, LASSO, RF, and SVM-RFE) was added to the final list of key lncRNAs. The genome-wide lncRNA expression profiles of 33 different types of cancers, a total of 9566 samples, available in The Cancer Genome Atlas (TCGA) were analyzed to discover the key lncRNAs. Our results showed that CAE performs better in feature selection, specially, in selecting small number of features, compared to LASSO, RF, and SVM-RFE. With the increasing number of selected features ranging from 10 to 500 lncRNAs, the accuracy of different feature selection approaches increases as - CAE: 70% to 96%; LASSO: 55% to 94%; RF: 38% to 95%; SVM-RFE: 50% to 94%. This study discovered a set of 69 lncRNAs that can identify the origin of 33 different cancers with an accuracy of 93%. Note that the accuracy could be higher using AE, which uses latent features for classification thus failing to correlate the origin of cancers with the actual features (lncRNAs).The proposed computational framework can be used as a diagnostic tool by the physicians to discover the origin of cancers using the expression profiles of lncRNAs. The discovered lncRNAs can be studied further by biologists or drug designer to identify possible targets for cancer therapy.
Abdullah Al Mamun 0004, Wenrui Duan, M. Mondal Ananda
BIBM3
2020 Deep Learning to Discover Cancer Glycome Genes Signifying the Origins of Cancer
abstract
Background: Aberrant protein glycosylation is a common feature of cancer and contributes to malignant behavior. However, how and to what extent the cellular glycome is involved in cancer development and progression is still undefined. The primary objective of this study is to conduct insilico identification of glycome genes that could reveal a signature of cancer using expression profiles of cancer genomes. There exists a list of ~500 glycome genes in several molecular categories. This study is based on the hypothesis that if the glycosylation is a common feature of cancer, there exists a shortlist of cancer glycome genes and their expression profiles should carry the signature capable of differentiating 33 different cancers available in The Cancer Genome Atlas (TCGA). Method: The distribution of cancer samples in TCGA is highly imbalanced, ranging from 36 for Cholangiocarcinoma (CHOL) to 1089 for Breast Cancer (BRCA). Supervised feature selection approaches to identify the signature genes would be biased to larger groups. We developed a computational framework using concrete autoencoder (CAE), a deep learning-based unsupervised feature selection algorithm, to find the cancer-related glycome genes. The criteria of optimal feature subset used in this study are (a) the number of features should be as few as possible, and (b) accuracy of classification using the selected features should be > 90%. Results: Our experiment showed a shortlist of glycome genes (132 genes) that can differentiate 33 different cancers with an accuracy of 92%. This study reflects that the cancer glycome genes signify the origins of cancer.
Abdullah Al Mamun 0004, Masrur Sobhan, Raihanul Bari Tanvir, Charles J. Dimitroff, M. Mondal Ananda
BIBM5
2020 Deep Learning to Discover Genomic Signatures for Racial Disparity in Lung Cancer
abstract
Background: In the United States, African American Males (AAM) have the highest lung cancer incidence and mortality rate compared to European American Males (EAM). Cigarette is considered the major risk factor for lung cancer, but smoking alone fails to interpret the rationale for developing lung cancer between AAM and EAM. The higher rates of lung cancer among AAM occur even though they have lower smoking rates, smoke fewer cigarettes per day, and are less likely to be heavy smokers than EAM. Identifying genomic signatures such as key genes that can differentiate lung cancers between AAM and EAM will be a stepping stone to comprehend the disparity of lung cancer between AAM and EAM.Method: The gene expression profiles of whole blood samples from AAM and EAM patients were used to identify the key genes that can differentiate the lung cancers between AAM and EAM. Due to the US population's imbalanced nature between AAM and EAM, the distribution of samples for the present study is also highly imbalanced (AAM: 15 and EAM: 153). Here, we developed a computational framework using a deep learning-based unsupervised feature selection approach, concrete autoencoder (CAE), which can select actual features rather than latent features. First, we showed that features such as differentially expressed genes (DEGs) discovered by a supervised statistical approach LIMMA could not differentiate lung cancers between AAM and EAM. Then we showed that the CAE could isolate essential features capable of differentiating lung cancers between AAM and EAM.Results: The proposed framework using CAE was able to detect 34 key features/genes, which outperforms all sets of DEGs identified using three different thresholds on fold change. Using the selected 34 genes, the Random Forest classifier was able to classify lung cancers among AAM and EAM with 99% accuracy and only one false negative.Conclusion: The proposed framework using CAE reveals the key genes that can differentiate lung tumors between AAM and EAM. These key genes can be used as biomarkers to understand the difference in lung cancer development between AAM and EAM. This study also showed that the CAE is capable of extracting relevant features from a highly imbalanced dataset.
Masrur Sobhan, Abdullah Al Mamun 0004, Raihanul Bari Tanvir, Mario Jacas Alfonso, M. Mondal Ananda
BIBM6
2020 Stage-Specific Co-expression Network Analysis for Cancer Biomarker Discovery
abstract
Identification of conserved gene network modules in different stages of cancer may lead to uncovering mechanisms behind cancer initiation and progression. This work is based on two hypotheses. Hypothesis-l: the network modules conserved in all cancer stages are potential biomarkers related to the trajectory of cancer development or progression of cancer from initiation to stage-to-stage to metastasis. Hypothesis-2: The network modules from a stage, which are not conserved in other stages, can be considered as the stage-specific biomarkers for diagnosis.To test the hypotheses, gene expression and clinical data of Breast Invasive Carcinoma (BRCA) from The Cancer Genome Atlas (TCGA) were used for analysis. Gene expression data was divided into five groups- stage I to stage IV and normal tissue samples. First, the co-expression networks for each of the four stages and normal samples were generated. Second, the modules from each of the stage-specific networks were discovered using weighted gene co-expression network analysis (WGCNA). Third, survival analysis was performed to identify the prognostically significant modules. Fourth, module preservation analysis was performed to determine whether a module from one stage is preserved in other cancer stages as well as in normal stage. Finally, gene ontology and pathway enrichment analyses were performed for the prognostically significant and conserved modules.The present study discovered several gene-network modules for breast cancer preserved in all cancer stages and are significant in overall survival; hence, they can be considered potential biomarkers for cancers, related to the trajectory of cancer development. The modules that were found not to be conserved in different stages can be considered as stage-specific biomarkers.
Raihanul Bari Tanvir, M. Mondal Ananda
BIBM2
2020 Computational identification of biomarker genes for lung cancer considering treatment and non-treatment studies
abstract
BACKGROUND: Lung cancer is the number one cancer killer in the world with more than 142,670 deaths estimated in the United States alone in the year 2019. Consequently, there is an overreaching need to identify the key biomarkers for lung cancer. The aim of this study is to computationally identify biomarker genes for lung cancer that can aid in its diagnosis and treatment. The gene expression profiles of two different types of studies, namely non-treatment and treatment, are considered for discovering biomarker genes. In non-treatment studies healthy samples are control and cancer samples are cases. Whereas, in treatment studies, controls are cancer cell lines without treatment and cases are cancer cell lines with treatment. RESULTS: The Differentially Expressed Genes (DEGs) for lung cancer were isolated from Gene Expression Omnibus (GEO) database using R software tool GEO2R. A total of 407 DEGs (254 upregulated and 153 downregulated) from non-treatment studies and 547 DEGs (133 upregulated and 414 downregulated) from treatment studies were isolated. Two Cytoscape apps, namely, CytoHubba and MCODE, were used for identifying biomarker genes from functional networks developed using DEG genes. This study discovered two distinct sets of biomarker genes - one from non-treatment studies and the other from treatment studies, each set containing 16 genes. Survival analysis results show that most non-treatment biomarker genes have prognostic capability by indicating low-expression groups have higher chance of survival compare to high-expression groups. Whereas, most treatment biomarkers have prognostic capability by indicating high-expression groups have higher chance of survival compare to low-expression groups. CONCLUSION: A computational framework is developed to identify biomarker genes for lung cancer using gene expression profiles. Two different types of studies - non-treatment and treatment - are considered for experiment. Most of the biomarker genes from non-treatment studies are part of mitosis and play vital role in DNA repair and cell-cycle regulation. Whereas, most of the biomarker genes from treatment studies are associated to ubiquitination and cellular response to stress. This study discovered a list of biomarkers, which would help experimental scientists to design a lab experiment for further exploration of detail dynamics of lung cancer development.
Mona Maharjan, Raihanul Bari Tanvir, Kamal Chowdhury, Wenrui Duan, M. Mondal Ananda
BMC Bioinform.5
2019 Pseudotime Based Discovery of Breast Cancer Heterogeneity
abstract
Breast cancer is highly sporadic and heterogeneous in nature. Even the patients with same clinical stage do not cluster together in terms of genomic profiles such as mRNA expression. In order to prevent and cure breast cancer completely, it is essential to decipher the detailed heterogeneity of breast cancer at genomic level. Putting the cancer patients on a time scale, which represents the trajectory of cancer development, may help discover the detailed heterogeneity. This in turn would help establish the mechanisms for prevention and complete cure of breast cancer. The goal of this study is to discover the heterogeneity of breast cancer by ordering the cancer patients using pseudotime. This is achieved through two objectives: First, a computational framework is developed to place the cancer patients on a time scale, meaning construct a trajectory of cancer development, by inferring pseudotime from static mRNA expression data; Second, discovering breast cancer heterogeneity at different time periods of the trajectory using statistical and machine learning techniques. In this study, the trajectory of breast cancer progression was constructed using static mRNA expression profiles of 1072 breast cancer patients by inferring pseudotime. Three sets of key genes discovered using supervised machine learning techniques are used to develop the trajectories. The first set of genes are PAM50 genes which is available in literature. The second and third sets of genes were discovered in the present study using the clinical stages of breast cancer (Stage-I, Stage-II, Stage-III, and Stage-IV). The proposed computational framework has the capability of deciphering heterogeneity in breast cancer at a granular level. The results also show the existence of multiple parallel trajectories at different time periods of cancer development or progression.
Tasmia Aqila, Abdullah Al Mamun 0004, M. Mondal Ananda
BIBM3
2019 Feature Selection and Classification Reveal Key lncRNAs for Multiple Cancers
abstract
Long noncoding RNA (lncRNA) plays key roles in tumorigenesis. Misexpression of lncRNA can lead to changes in expression profiles of various target genes, which are involved in cancer initiation and progression. So, identifying key lncRNAs for a cancer would help develop the cancer therapy. Usually, to identify key lncRNAs for a cancer, expression profiles of lncRNAs for normal and cancer samples are required. But, this kind of data are not available for all cancers. In the present study, a computational framework is developed to identify cancer specific key lncRNAs using the lncRNA expression of cancer patients only. The framework consists of two state-of-the-art feature selection techniques - Recursive Feature Elimination (RFE) and Least Absolute Shrinkage and Selection Operator (LASSO); and five machine learning models - Naive Bayes, K-Nearest Neighbor, Random Forest, Support Vector Machine, and Deep Neural Network. For experiment, expression values of lncRNAs for 8 cancers - BLCA, CESC, COAD, HNSC, KIRP, LGG, LIHC, and LUAD - from TCGA are used. The combined dataset consists of 3,656 patients with expression values of 12,309 lncRNAs. Important features or key lncRNAs are identified by using feature selection algorithms RFE and LASSO. Capability of these key lncRNAs in classifying 8 different cancers is checked by the performance of five classification models. This study identified 37 key lncRNAs that can classify 8 different cancer types with an accuracy ranging from 94% to 97%. Finally, survival analysis supports that the discovered key lncRNAs are capable of differentiating between high-risk and low-risk patients.
Abdullah Al Mamun 0004, M. Mondal Ananda
BIBM2
2019 Cancer Biomarker Discovery from Gene Co-expression Networks Using Community Detection Methods
abstract
Finding the network biomarkers of cancers and the analysis of cancer driving genes that are involved in these biomarkers are essential for understanding the dynamics of cancer. Clusters of genes in co-expression networks are commonly known as functional units. This work is based on the hypothesis that the dense clusters or communities in the gene co-expression networks of cancer patients may represent functional units regarding cancer initiation and progression. In this study, RNA-seq gene expression data of three cancers - Breast Invasive Carcinoma (BRCA), Colorectal Adenocarcinoma (COAD) and Glioblastoma Multiforme (GBM) - from The Cancer Genome Atlas (TCGA) are used to construct gene co-expression networks using Pearson Correlation. Six well-known community detection algorithms are applied on these networks to identify communities with five or more genes. A permutation test is performed to further mine the communities that are conserved in other cancers, thus calling them conserved communities. Then survival analysis is performed on clinical data of three cancers using the conserved community genes as prognostic co-variates. The communities that could distinguish the cancer patients between high- and low-risk groups are considered as cancer biomarkers. In the present study, 16 such network biomarkers are discovered.
Raihanul Bari Tanvir, M. Mondal Ananda
BIBM2
2019 Texture-based Deep Learning for Effective Histopathological Cancer Image Classification
abstract
Automatic histopathological Whole Slide Image (WSI) analysis for cancer classification has been highlighted along with the advancements in microscopic imaging techniques, since manual examination and diagnosis with WSIs are time- and cost-consuming. Recently, deep convolutional neural networks have succeeded in histopathological image analysis. However, despite the success of the development, there are still opportunities for further enhancements. In this paper, we propose a novel cancer texture-based deep neural network (CAT-Net) that learns scalable morphological features from histopathological WSIs. The innovation of CAT-Net is twofold: (1) capturing invariant spatial patterns by dilated convolutional layers and (2) improving predictive performance while reducing model complexity. Moreover, CAT-Net can provide discriminative morphological (texture) patterns formed on cancerous regions of histopathological images comparing to normal regions. We elucidated how our proposed method, CAT-Net, captures morphological patterns of interest in hierarchical levels in the model. The proposed method out-performed the current state-of-the-art benchmark methods on accuracy, precision, recall, and F1 score.
Nelson Zange Tsaku, Sai Kosaraju, Tasmia Aqila, Mohammad Masum, Dae Hyun Song, M. Mondal Ananda, Hyun Min Koh, Mingon Kang
BIBM6
2018 Graph Theoretic Concepts as the Building Blocks for Disease Initiation and Progression at Protein Network Level: Identification and Challenges
M. Mondal Ananda, Cornelia Ada Schultz, Markea Sheppard, Jasmine Carson, Raihanul Bari Tanvir, Tasmia Aqila
BIBM1
2015 Diffusion kernel to identify missing PPIs in protein network biomarker
abstract
Little focus has been placed on neighborhood proteins in the protein-protein interaction (PPI) network that do not physically interact with each other but have a higher likelihood to interact than the actual PPIs in the network. Identifying these missing PPIs would complete the protein network biomarker representing a disease. In the present study, we check the capability of diffusion kernel in identifying these missing PPIs. A diffusion kernel is a computational framework which is based on a physical phenomenon of gas diffusion and a computer science concept of random walk on a graph or network. We seek to predict probable missing PPIs related to Allergy and Asthma using diffusion kernel, by employing a threshold on the kernel values. The completed protein network biomarker can be used as a better predictor for disease identification and classification. This would also help in better understanding of disease mechanism at protein network level. Our results show that the network with high PPI score has better accuracy of predicting missing PPIs.
Dominic K. Bett, M. Mondal Ananda
BIBM2
2014 STRING PPI Score to Characterize Protein Subnetwork Biomarkers for Human Diseases and Pathways
abstract
Protein sub network biomarkers for 144 diseases and pathways are analyzed in terms of protein-protein interaction (PPI) score available in STRING database. Most of the sub network biomarker (SNB) studies are to classify disease samples from the control. But no de novo algorithm is available to identify SNB from the whole genome PPI network without the knowledge of differentially expressed genes. Recently, based on mouse model, researchers showed that there exists a dynamical network biomarker which can distinguish among the normal state, pre-disease state, and disease state of a disease progression. But, most of the gene expression data for human diseases are at the disease state. No data is available for the first two stages of a disease. Understanding the network behavior of a disease at disease state might help in the development of de novo algorithm for predicting protein SNBs not only for disease state but also for early stages of a disease or early warning signals. PPI score in STRING database represents a rough estimate of how likely a given interaction describes a functional linkage between two proteins. So, analyzing protein SNB for human diseases at disease state with respect to PPI score may shed some light in the development of de novo models for predicting SNB. A simple brute force approach is used to isolate the SNB for a disease or pathway from the genome-wide PPI network by projecting the corresponding differentially expressed proteins. Then the SNBs are analyzed in terms of PPI score. Our investigation shows that higher is the PPI score of a network is more likely to produce a true SNB for a disease. Results also show that Physical PPIs with high score are more capable of producing a true SNB.
Prayas Timalsina, Kevin Charles, M. Mondal Ananda
BIBE3
2012 Minimalist ensemble algorithms for genome-wide protein localization prediction
abstract
BACKGROUND: Computational prediction of protein subcellular localization can greatly help to elucidate its functions. Despite the existence of dozens of protein localization prediction algorithms, the prediction accuracy and coverage are still low. Several ensemble algorithms have been proposed to improve the prediction performance, which usually include as many as 10 or more individual localization algorithms. However, their performance is still limited by the running complexity and redundancy among individual prediction algorithms. RESULTS: This paper proposed a novel method for rational design of minimalist ensemble algorithms for practical genome-wide protein subcellular localization prediction. The algorithm is based on combining a feature selection based filter and a logistic regression classifier. Using a novel concept of contribution scores, we analyzed issues of algorithm redundancy, consensus mistakes, and algorithm complementarity in designing ensemble algorithms. We applied the proposed minimalist logistic regression (LR) ensemble algorithm to two genome-wide datasets of Yeast and Human and compared its performance with current ensemble algorithms. Experimental results showed that the minimalist ensemble algorithm can achieve high prediction accuracy with only 1/3 to 1/2 of individual predictors of current ensemble algorithms, which greatly reduces computational complexity and running time. It was found that the high performance ensemble algorithms are usually composed of the predictors that together cover most of available features. Compared to the best individual predictor, our ensemble algorithm improved the prediction accuracy from AUC score of 0.558 to 0.707 for the Yeast dataset and from 0.628 to 0.646 for the Human dataset. Compared with popular weighted voting based ensemble algorithms, our classifier-based ensemble algorithms achieved much better performance without suffering from inclusion of too many individual predictors. CONCLUSIONS: We proposed a method for rational design of minimalist ensemble algorithms using feature selection and classifiers. The proposed minimalist ensemble algorithm based on logistic regression can achieve equal or better prediction performance while using only half or one-third of individual predictors compared to other ensemble algorithms. The results also suggested that meta-predictors that take advantage of a variety of features by combining individual predictors tend to achieve the best performance. The LR ensemble server and related benchmark datasets are available at http://mleg.cse.sc.edu/LRensemble/cgi-bin/predict.cgi.
Jhih-rong Lin, M. Mondal Ananda, Jianjun Hu
BMC Bioinform.2
2010 NetLoc: Network based protein localization prediction using protein-protein interaction and co-expression networks
abstract
Recent studies showed that protein-protein interaction network based features can significantly improve the prediction of protein subcellular localization. However, it is unclear whether network prediction models or other types of protein-protein correlation networks would also improve localization prediction. We present NetLoc, a novel diffusion kernel-based logistic regression (KLR) algorithm for predicting protein subcellular localization using four types of protein networks including physical protein-protein interaction (PPPI) networks, genetic PPI networks (GPPI), mixed PPI networks (MPPI), and co-expression networks (COEXP). We applied NetLoc to yeast protein localization prediction. The results showed that protein networks can provide rich information for protein localization prediction, achieving prediction performance up to AUC score of 0.93. We also showed that networks with high connectivity and high percentage of interacting protein pairs targeting the same location lead to better prediction performance. We found that physical PPPI is better than GPPI which is better than COEXP in terms of localization prediction. The prediction performance (AUC) using the yeast PPPI network ranges between 0.71 and 0.93 for 7 locations. Compared to the previous network feature based prediction algorithm which achieved AUC scores of (0.49 and 0.52) on the yeast PPI network of the DIP database, NetLoc achieved significantly better overall performance with the AUC of 0.74.
M. Mondal Ananda, Jianjun Hu
BIBM1