EDBT 2026 Demo / reviewers in the wild / expert
Kim-Anh Do
dblp:99/8654
· DBLP profile ↗
20ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 9 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CAT: a conditional association test for microbiome data using a permutation approachabstractIn microbiome analysis, researchers often seek to identify taxonomic features associated with an outcome of interest. However, microbiome features are intercorrelated and linked by phylogenetic relationships, making it challenging to assess the association between an individual feature and an outcome. This paper proposes a novel conditional association test, CAT, that can account for other features and phylogenetic relatedness when testing the association between a feature and an outcome. CAT adopts a permutation approach, measuring the importance of a feature in predicting the outcome by permuting operational taxonomic unit/amplicon sequence variant counts belonging to that feature from the data and quantifying how much the association with the outcome is weakened through the change in the coefficient of determination $R^{2}$. Compared with marginal association tests, it focuses on the added value of a feature in explaining outcome variation that is not captured by other features. By leveraging global tests including PERMANOVA and MiRKAT-based methods, CAT allows association testing for continuous, binary, categorical, count, survival, and correlated outcomes. We demonstrate through simulation studies that CAT can provide a direct quantification of feature importance that is distinct from that of marginal association tests, and illustrate CAT with applications to two real-world studies on the microbiome in melanoma patients: one examining the role of the microbiome in shaping immunotherapy response, and one investigating the association between the microbiome and survival outcomes. Our results illustrate the potential of CAT to inform the design of microbiome interventions aimed at improving clinical outcomes. Yushu Shi, Kim-Anh Do, Robert R. Jenq, Christine B. Peterson |
Briefings Bioinform. | 3 |
| 2025 | Rank-based learning: a novel high-throughput algorithm resilient to missing data and effective for datasets with small sample sizeabstractHigh-throughput omics data present challenges for binary classification due to platform variability, batch effects, missing values, and high dimensionality. This study presents a novel Rank-Based Learning (RBL) method that leverages relative feature rankings to improve robustness and generalizability. We evaluated RBL against established methods like Logistic Regression (LR) and Random Forest (RF) using simulated data and two real-world plasma proteomics datasets: early-stage small cell lung cancer (SCLC) and duodenopancreatic neuroendocrine tumors (dpNET) in patients with Multiple Endocrine Neoplasia type 1 (MEN1). In simulation experiments, RBL outperformed LR under conditions involving batch effects, missing data, and varying numbers of true differential features. In SCLC, RBL yielded a test AUC of 0.76 (95% CI: 0.42-1.00), surpassing LR with Lasso (0.65 [95% CI: 0.47-0.84]) and RF with feature importance (0.59 [95% CI: 0.33-0.87]). In dpNET, RBL achieved an AUC of 0.83 (95% CI: 0.67-0.97) on the development set and 0.80 (95% CI: 0.54-0.98) on the test set, outperforming LR with Lasso (0.57 [95% CI: 0.40-0.77]) and RF with feature importance (0.53 [95% CI: 0.29-0.77]). By emphasizing feature ranking rather than absolute expression levels, RBL effectively mitigates the impact of non-biological variation. Overall, RBL improves the predictive accuracy of diagnostic models for complex diseases and provides a promising framework for developing more reliable and generalizable diagnostic tools from omics data, moving them closer to clinical application. Lulu Song, Hamid Khoshfekr Rudsari, Johannes F. Fahrmann, Jody V. Vykoukal, Sam Hanash, James P. Long, Kim-Anh Do, Ehsan Irajizad |
Briefings Bioinform. | 7 |
| 2025 | AUPRC: a metric for evaluating the performance of in-silico perturbation methods in identifying differentially expressed genesabstractIn silico perturbation models, computational methods that can predict cellular responses to perturbations, present an opportunity to reduce the need for costly and time-intensive in vitro experiments. Many recently proposed models predict high-dimensional cellular responses, such as gene or protein expression to perturbations such as gene knockout or drugs. However, evaluating in silico performance has largely relied on metrics such as $R^{2}$, which assess overall prediction accuracy but fail to capture biologically significant outcomes like the identification of differentially expressed (DE) genes. In this study, we present a novel evaluation framework that introduces the AUPRC metric to assess the precision and recall of DE gene predictions. By applying this framework to both single-cell and pseudo-bulked datasets, we systematically benchmark simple and advanced computational models. Our results highlight a significant discrepancy between $R^{2}$ and AUPRC, with models achieving high $R^{2}$ values but struggling to identify DE genes, as reflected in their low AUPRC values. This finding underscores the limitations of traditional evaluation metrics and the importance of biologically relevant assessments. Our framework provides a more comprehensive understanding of model capabilities, advancing the application of computational approaches in cellular perturbation research. Hongxu Zhu, Amir Asiaee, Leila Azinfar, Jun Li 0068, Ehsan Irajizad, Kim-Anh Do, James P. Long |
Briefings Bioinform. | 7 |
| 2025 | Causal models and prediction in cell line perturbation experimentsabstractIn cell line perturbation experiments, a collection of cells is perturbed with external agents and responses such as protein expression measured. Due to cost constraints, only a small fraction of all possible perturbations can be tested in vitro. This has led to the development of computational models that can predict cellular responses to perturbations in silico. A central challenge for these models is to predict the effect of new, previously untested perturbations that were not used in the training data. Here we propose causal structural equations for modeling how perturbations effect cells. From this model, we derive two estimators for predicting responses: a Linear Regression (LR) estimator and a causal structure learning estimator that we term Causal Structure Regression (CSR). The CSR estimator requires more assumptions than LR, but can predict the effects of drugs that were not applied in the training data. Next we present Cellbox, a recently proposed system of ordinary differential equations (ODEs) based model that obtained the best prediction performance on a Melanoma cell line perturbation data set (Yuan et al. in Cell Syst 12:128-140, 2021). We derive analytic results that show a close connection between CSR and Cellbox, providing a new causal interpretation for the Cellbox model. We compare LR and CSR/Cellbox in simulations, highlighting the strengths and weaknesses of the two approaches. Finally we compare the performance of LR and CSR/Cellbox on the benchmark Melanoma data set. We find that the LR model has comparable or slightly better performance than Cellbox. James P. Long, Shohei Shimizu, Thong Pham, Kim-Anh Do |
BMC Bioinform. | 5 |
| 2023 | Easy NanoString nCounter data analysis with the NanoTubeabstractSUMMARY: The NanoTube is an open-source pipeline that simplifies the processing, quality control, normalization and analysis of NanoString nCounter gene expression data. It is implemented in an extensible R library, which performs a variety of gene expression analysis techniques and contains additional functions for integration with other R libraries performing advanced NanoString analysis techniques. Additionally, the NanoTube web application is available as a simple tool for researchers without programming expertise. AVAILABILITY AND IMPLEMENTATION: The NanoTube R package is available on Bioconductor under the GPL-3 license (https://www.bioconductor.org/packages/NanoTube/). The R-Shiny application can be downloaded at https://github.com/calebclass/Shiny-NanoTube, or a simplified version of this application can be run on all major browsers, at https://research.butler.edu/nanotube/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Caleb A. Class, Caiden J. Lukan, Christopher A. Bristow, Kim-Anh Do |
Bioinform. | 4 |
| 2023 | A unified mediation analysis framework for integrative cancer proteogenomics with clinical outcomesabstractMOTIVATION: Multilevel molecular profiling of tumors and the integrative analysis with clinical outcomes have enabled a deeper characterization of cancer treatment. Mediation analysis has emerged as a promising statistical tool to identify and quantify the intermediate mechanisms by which a gene affects an outcome. However, existing methods lack a unified approach to handle various types of outcome variables, making them unsuitable for high-throughput molecular profiling data with highly interconnected variables. RESULTS: We develop a general mediation analysis framework for proteogenomic data that include multiple exposures, multivariate mediators on various scales of effects as appropriate for continuous, binary and survival outcomes. Our estimation method avoids imposing constraints on model parameters such as the rare disease assumption, while accommodating multiple exposures and high-dimensional mediators. We compare our approach to other methods in extensive simulation studies at a range of sample sizes, disease prevalence and number of false mediators. Using kidney renal clear cell carcinoma proteogenomic data, we identify genes that are mediated by proteins and the underlying mechanisms on various survival outcomes that capture short- and long-term disease-specific clinical characteristics. AVAILABILITY AND IMPLEMENTATION: Software is made available in an R package (https://github.com/longjp/mediateR). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Licai Huang, James P. Long, Ehsan Irajizad, James D. Doecke, Kim-Anh Do, Min Jin Ha |
Bioinform. | 5 |
| 2022 | A machine learning-based method for automatically identifying novel cells in annotating single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) has been widely used to decompose complex tissues into functionally distinct cell types. The first and usually the most important step of scRNA-seq data analysis is to accurately annotate the cell labels. In recent years, many supervised annotation methods have been developed and shown to be more convenient and accurate than unsupervised cell clustering. One challenge faced by all the supervised annotation methods is the identification of the novel cell type, which is defined as the cell type that is not present in the training data, only exists in the testing data. Existing methods usually label the cells simply based on the correlation coefficients or confidence scores, which sometimes results in an excessive number of unlabeled cells. RESULTS: We developed a straightforward yet effective method combining autoencoder with iterative feature selection to automatically identify novel cells from scRNA-seq data. Our method trains an autoencoder with the labeled training data and applies the autoencoder to the testing data to obtain reconstruction errors. By iteratively selecting features that demonstrate a bi-modal pattern and reclustering the cells using the selected feature, our method can accurately identify novel cells that are not present in the training data. We further combined this approach with a support vector machine to provide a complete solution for annotating the full range of cell types. Extensive numerical experiments using five real scRNA-seq datasets demonstrated favorable performance of the proposed method over existing methods serving similar purposes. AVAILABILITY AND IMPLEMENTATION: Our R software package CAMLU is publicly available through the Zenodo repository (https://doi.org/10.5281/zenodo.7054422) or GitHub repository (https://github.com/ziyili20/CAMLU). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziyi Li 0001, Yizhuo Wang 0002, Irene Ganan-Gomez, Simona Colla, Kim-Anh Do |
Bioinform. | 5 |
| 2022 | CNGPLD: case-control copy-number analysis using Gaussian process latent differenceabstractMOTIVATION: Cross-sectional analyses of primary cancer genomes have identified regions of recurrent somatic copy-number alteration, many of which result from positive selection during cancer formation and contain driver genes. However, no effective approach exists for identifying genomic loci under significantly different degrees of selection in cancers of different subtypes, anatomic sites or disease stages. RESULTS: CNGPLD is a new tool for performing case-control somatic copy-number analysis that facilitates the discovery of differentially amplified or deleted copy-number aberrations in a case group of cancer compared with a control group of cancer. This tool uses a Gaussian process statistical framework in order to account for the covariance structure of copy-number data along genomic coordinates and to control the false discovery rate at the region level. AVAILABILITY AND IMPLEMENTATION: CNGPLD is freely available at https://bitbucket.org/djhshih/cngpld as an R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. David J. H. Shih, Ruoxing Li, W. Jim Zheng, Kim-Anh Do, Shiaw-Yih Lin, Scott L. Carter |
Bioinform. | 5 |
| 2021 | ProgPerm: Progressive permutation for a dynamic representation of the robustness of microbiome discoveriesabstractBACKGROUND: Identification of features is a critical task in microbiome studies that is complicated by the fact that microbial data are high dimensional and heterogeneous. Masked by the complexity of the data, the problem of separating signals (differential features between groups) from noise (features that are not differential between groups) becomes challenging and troublesome. For instance, when performing differential abundance tests, multiple testing adjustments tend to be overconservative, as the probability of a type I error (false positive) increases dramatically with the large numbers of hypotheses. Moreover, the grouping effect of interest can be obscured by heterogeneity. These factors can incorrectly lead to the conclusion that there are no differences in the microbiome compositions. RESULTS: We translate and represent the problem of identifying differential features, which are differential in two-group comparisons (e.g., treatment versus control), as a dynamic layout of separating the signal from its random background. More specifically, we progressively permute the grouping factor labels of the microbiome samples and perform multiple differential abundance tests in each scenario. We then compare the signal strength of the most differential features from the original data with their performance in permutations, and will observe a visually apparent decreasing trend if these features are true positives identified from the data. Simulations and applications on real data show that the proposed method creates a U-curve when plotting the number of significant features versus the proportion of mixing. The shape of the U-Curve can convey the strength of the overall association between the microbiome and the grouping factor. We also define a fragility index to measure the robustness of the discoveries. Finally, we recommend the identified features by comparing p-values in the observed data with p-values in the fully mixed data. CONCLUSIONS: We have developed this into a user-friendly and efficient R-shiny tool with visualizations. By default, we use the Wilcoxon rank sum test to compute the p-values, since it is a robust nonparametric test. Our proposed method can also utilize p-values obtained from other testing methods, such as DESeq. This demonstrates the potential of the progressive permutation method to be extended to new settings. Yushu Shi, Kim-Anh Do, Christine B. Peterson, Robert R. Jenq |
BMC Bioinform. | 3 |
| 2020 | NExUS: Bayesian simultaneous network estimation across unequal sample sizesabstractMOTIVATION: Network-based analyses of high-throughput genomics data provide a holistic, systems-level understanding of various biological mechanisms for a common population. However, when estimating multiple networks across heterogeneous sub-populations, varying sample sizes pose a challenge in the estimation and inference, as network differences may be driven by differences in power. We are particularly interested in addressing this challenge in the context of proteomic networks for related cancers, as the number of subjects available for rare cancer (sub-)types is often limited. RESULTS: We develop NExUS (Network Estimation across Unequal Sample sizes), a Bayesian method that enables joint learning of multiple networks while avoiding artefactual relationship between sample size and network sparsity. We demonstrate through simulations that NExUS outperforms existing network estimation methods in this context, and apply it to learn network similarity and shared pathway activity for groups of cancers with related origins represented in The Cancer Genome Atlas (TCGA) proteomic data. AVAILABILITY AND IMPLEMENTATION: The NExUS source code is freely available for download at https://github.com/priyamdas2/NExUS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Priyam Das, Christine B. Peterson, Kim-Anh Do, Rehan Akbani, Veerabhadran Baladandayuthapani |
Bioinform. | 3 |
| 2020 | aPCoA: covariate adjusted principal coordinates analysisabstractSUMMARY: In fields, such as ecology, microbiology and genomics, non-Euclidean distances are widely applied to describe pairwise dissimilarity between samples. Given these pairwise distances, principal coordinates analysis is commonly used to construct a visualization of the data. However, confounding covariates can make patterns related to the scientific question of interest difficult to observe. We provide adjusted principal coordinates analysis as an easy-to-use tool, available as both an R package and a Shiny app, to improve data visualization in this context, enabling enhanced presentation of the effects of interest. AVAILABILITY AND IMPLEMENTATION: The R package 'aPCoA' and Shiny app can be accessed at https://cran.r-project.org/web/packages/aPCoA/index.html and https://biostatistics.mdanderson.org/shinyapps/aPCoA/. Yushu Shi, Kim-Anh Do, Christine B. Peterson, Robert R. Jenq |
Bioinform. | 3 |
| 2020 | Compositional zero-inflated network estimation for microbiome dataabstractBACKGROUND: The estimation of microbial networks can provide important insight into the ecological relationships among the organisms that comprise the microbiome. However, there are a number of critical statistical challenges in the inference of such networks from high-throughput data. Since the abundances in each sample are constrained to have a fixed sum and there is incomplete overlap in microbial populations across subjects, the data are both compositional and zero-inflated. RESULTS: We propose the COmpositional Zero-Inflated Network Estimation (COZINE) method for inference of microbial networks which addresses these critical aspects of the data while maintaining computational scalability. COZINE relies on the multivariate Hurdle model to infer a sparse set of conditional dependencies which reflect not only relationships among the continuous values, but also among binary indicators of presence or absence and between the binary and continuous representations of the data. Our simulation results show that the proposed method is better able to capture various types of microbial relationships than existing approaches. We demonstrate the utility of the method with an application to understanding the oral microbiome network in a cohort of leukemic patients. CONCLUSIONS: Our proposed method addresses important challenges in microbiome network estimation, and can be effectively applied to discover various types of dependence relationships in microbial communities. The procedure we have developed, which we refer to as COZINE, is available online at https://github.com/MinJinHa/COZINE . Min Jin Ha, Junghi Kim, Jessica Galloway-Pena, Kim-Anh Do, Christine B. Peterson |
BMC Bioinform. | 4 |
| 2018 | iDINGO - integrative differential network analysis in genomics with Shiny applicationabstractMotivation: Differential network analysis is an important way to understand network rewiring involved in disease progression and development. Building differential networks from multiple 'omics data provides insight into the holistic differences of the interactive system under different patient-specific groups. DINGO was developed to infer group-specific dependencies and build differential networks. However, DINGO and other existing tools are limited to analyze data arising from a single platform, and modeling each of the multiple 'omics data independently does not account for the hierarchical structure of the data. Results: We developed the iDINGO R package to estimate group-specific dependencies and make inferences on the integrative differential networks, considering the biological hierarchy among the platforms. A Shiny application has also been developed to facilitate easier analysis and visualization of results, including integrative differential networks and hub gene identification across platforms. Availability and implementation: R package is available on CRAN (https://cran.r-project.org/web/packages/iDINGO) and Shiny application at https://github.com/MinJinHa/iDINGO. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Caleb A. Class, Min Jin Ha, Veerabhadran Baladandayuthapani, Kim-Anh Do |
Bioinform. | 4 |
| 2016 | SIFORM: shared informative factor models for integration of multi-platform bioinformatic dataabstractMOTIVATION: High-dimensional omic data derived from different technological platforms have been extensively used to facilitate comprehensive understanding of disease mechanisms and to determine personalized health treatments. Numerous studies have integrated multi-platform omic data; however, few have efficiently and simultaneously addressed the problems that arise from high dimensionality and complex correlations. RESULTS: We propose a statistical framework of shared informative factor models that can jointly analyze multi-platform omic data and explore their associations with a disease phenotype. The common disease-associated sample characteristics across different data types can be captured through the shared structure space, while the corresponding weights of genetic variables directly index the strengths of their association with the phenotype. Extensive simulation studies demonstrate the performance of the proposed method in terms of biomarker detection accuracy via comparisons with three popular regularized regression methods. We also apply the proposed method to The Cancer Genome Atlas lung adenocarcinoma dataset to jointly explore associations of mRNA expression and protein expression with smoking status. Many of the identified biomarkers belong to key pathways for lung tumorigenesis, some of which are known to show differential expression across smoking levels. We discover potential biomarkers that reveal different mechanisms of lung tumorigenesis between light smokers and heavy smokers. AVAILABILITY AND IMPLEMENTATION: R code to implement the new method can be downloaded from http://odin.mdacc.tmc.edu/jhhu/ CONTACT: [email protected]. Xuebei An, Kim-Anh Do |
Bioinform. | 3 |
| 2015 | DINGO: differential network analysis in genomicsabstractMOTIVATION: Cancer progression and development are initiated by aberrations in various molecular networks through coordinated changes across multiple genes and pathways. It is important to understand how these networks change under different stress conditions and/or patient-specific groups to infer differential patterns of activation and inhibition. Existing methods are limited to correlation networks that are independently estimated from separate group-specific data and without due consideration of relationships that are conserved across multiple groups. METHOD: We propose a pathway-based differential network analysis in genomics (DINGO) model for estimating group-specific networks and making inference on the differential networks. DINGO jointly estimates the group-specific conditional dependencies by decomposing them into global and group-specific components. The delineation of these components allows for a more refined picture of the major driver and passenger events in the elucidation of cancer progression and development. RESULTS: Simulation studies demonstrate that DINGO provides more accurate group-specific conditional dependencies than achieved by using separate estimation approaches. We apply DINGO to key signaling pathways in glioblastoma to build differential networks for long-term survivors and short-term survivors in The Cancer Genome Atlas. The hub genes found by mRNA expression, DNA copy number, methylation and microRNA expression reveal several important roles in glioblastoma progression. AVAILABILITY AND IMPLEMENTATION: R Package at: odin.mdacc.tmc.edu/∼vbaladan. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Min Jin Ha, Veerabhadran Baladandayuthapani, Kim-Anh Do |
Bioinform. | 3 |
| 2013 | iBAG: integrative Bayesian analysis of high-dimensional multiplatform genomics dataabstractMOTIVATION: Analyzing data from multi-platform genomics experiments combined with patients' clinical outcomes helps us understand the complex biological processes that characterize a disease, as well as how these processes relate to the development of the disease. Current data integration approaches are limited in that they do not consider the fundamental biological relationships that exist among the data obtained from different platforms. Statistical Model: We propose an integrative Bayesian analysis of genomics data (iBAG) framework for identifying important genes/biomarkers that are associated with clinical outcome. This framework uses hierarchical modeling to combine the data obtained from multiple platforms into one model. RESULTS: We assess the performance of our methods using several synthetic and real examples. Simulations show our integrative methods to have higher power to detect disease-related genes than non-integrative methods. Using the Cancer Genome Atlas glioblastoma dataset, we apply the iBAG model to integrate gene expression and methylation data to study their associations with patient survival. Our proposed method discovers multiple methylation-regulated genes that are related to patient survival, most of which have important biological functions in other diseases but have not been previously studied in glioblastoma. AVAILABILITY: http://odin.mdacc.tmc.edu/∼vbaladan/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Veerabhadran Baladandayuthapani, Jeffrey S. Morris, Bradley M. Broom, Ganiraju Manyam, Kim-Anh Do |
Bioinform. | 6 |
| 2013 | Integrative network-based Bayesian analysis of diverse genomics dataabstractBACKGROUND: In order to better understand cancer as a complex disease with multiple genetic and epigenetic factors, it is vital to model the fundamental biological relationships among these alterations as well as their relationships with important clinical outcomes. METHODS: We develop an integrative network-based Bayesian analysis (iNET) approach that allows us to jointly analyze multi-platform high-dimensional genomic data in a computationally efficient manner. The iNET approach is formulated as an objective Bayesian model selection problem for Gaussian graphical models to model joint dependencies among platform-specific features using known biological mechanisms. Using both simulated datasets and a glioblastoma (GBM) study from The Cancer Genome Atlas (TCGA), we illustrate the iNET approach via integrating three data types, microRNA, gene expression (mRNA), and patient survival time. RESULTS: We show that the iNET approach has greater power in identifying cancer-related microRNAs than non-integrative approaches based on realistic simulated datasets. In the TCGA GBM study, we found many mRNA-microRNA pairs and microRNAs that are associated with patient survival time, with some of these associations identified in previous studies. CONCLUSIONS: The iNET discovers relationships consistent with the underlying biological mechanisms among these variables, as well as identifying important biomarkers that are potentially relevant to patient survival. In addition, we identified some microRNAs that can potentially affect patient survival which are missed by non-integrative approaches. Veerabhadran Baladandayuthapani, Christopher C. Holmes, Kim-Anh Do |
BMC Bioinform. | 4 |
| 2012 | Model averaging strategies for structure learning in Bayesian networks with limited dataabstractBACKGROUND: Considerable progress has been made on algorithms for learning the structure of Bayesian networks from data. Model averaging by using bootstrap replicates with feature selection by thresholding is a widely used solution for learning features with high confidence. Yet, in the context of limited data many questions remain unanswered. What scoring functions are most effective for model averaging? Does the bias arising from the discreteness of the bootstrap significantly affect learning performance? Is it better to pick the single best network or to average multiple networks learnt from each bootstrap resample? How should thresholds for learning statistically significant features be selected? RESULTS: The best scoring functions are Dirichlet Prior Scoring Metric with small λ and the Bayesian Dirichlet metric. Correcting the bias arising from the discreteness of the bootstrap worsens learning performance. It is better to pick the single best network learnt from each bootstrap resample. We describe a permutation based method for determining significance thresholds for feature selection in bagged models. We show that in contexts with limited data, Bayesian bagging using the Dirichlet Prior Scoring Metric (DPSM) is the most effective learning strategy, and that modifying the scoring function to penalize complex networks hampers model averaging. We establish these results using a systematic study of two well-known benchmarks, specifically ALARM and INSURANCE. We also apply our network construction method to gene expression data from the Cancer Genome Atlas Glioblastoma multiforme dataset and show that survival is related to clinical covariates age and gender and clusters for interferon induced genes and growth inhibition genes. CONCLUSIONS: For small data sets, our approach performs significantly better than previously published methods. Bradley M. Broom, Kim-Anh Do, Devika Subramanian |
BMC Bioinform. | 2 |
| 2011 | Bayesian ensemble methods for survival prediction in gene expression dataabstractMOTIVATION: We propose a Bayesian ensemble method for survival prediction in high-dimensional gene expression data. We specify a fully Bayesian hierarchical approach based on an ensemble 'sum-of-trees' model and illustrate our method using three popular survival models. Our non-parametric method incorporates both additive and interaction effects between genes, which results in high predictive accuracy compared with other methods. In addition, our method provides model-free variable selection of important prognostic markers based on controlling the false discovery rates; thus providing a unified procedure to select relevant genes and predict survivor functions. RESULTS: We assess the performance of our method several simulated and real microarray datasets. We show that our method selects genes potentially related to the development of the disease as well as yields predictive performance that is very competitive to many other existing methods. AVAILABILITY: http://works.bepress.com/veera/1/. Vinícius Bonato, Veerabhadran Baladandayuthapani, Bradley M. Broom, Erik P. Sulman, Kenneth D. Aldape, Kim-Anh Do |
Bioinform. | 6 |
| 1992 | Alternative preprocessing techniques for discrete hidden Markov model phoneme recognition
Andrew Tridgell, J. Bruce Millar, Kim-Anh Do |
ICSLP | 3 |