Kevin R. Coombes

dblp:75/366 · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
5since 2021 · last 2022
0000-0002-7630-2123ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 5 since 2021
YearPublicationVenuePosition
2022 CEDA: integrating gene expression data with CRISPR-pooled screen data identifies essential genes with higher expression
abstract
MOTIVATION: Clustered regularly interspaced short palindromic repeats (CRISPR)-based genetic perturbation screen is a powerful tool to probe gene function. However, experimental noises, especially for the lowly expressed genes, need to be accounted for to maintain proper control of false positive rate. METHODS: We develop a statistical method, named CRISPR screen with Expression Data Analysis (CEDA), to integrate gene expression profiles and CRISPR screen data for identifying essential genes. CEDA stratifies genes based on expression level and adopts a three-component mixture model for the log-fold change of single-guide RNAs (sgRNAs). Empirical Bayesian prior and expectation-maximization algorithm are used for parameter estimation and false discovery rate inference. RESULTS: Taking advantage of gene expression data, CEDA identifies essential genes with higher expression. Compared to existing methods, CEDA shows comparable reliability but higher sensitivity in detecting essential genes with moderate sgRNA fold change. Therefore, using the same CRISPR data, CEDA generates an additional hit gene list. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lianbo Yu, Haoran Li 0015, Kevin R. Coombes, Kin Fai Au, Lijun Cheng, Lang Li 0001
Bioinform.5
2021 Mercator: a pipeline for multi-method, unsupervised visualization and distance generation
abstract
SUMMARY: Unsupervised machine learning provides tools for researchers to uncover latent patterns in large-scale data, based on calculated distances between observations. Methods to visualize high-dimensional data based on these distances can elucidate subtypes and interactions within multi-dimensional and high-throughput data. However, researchers can select from a vast number of distance metrics and visualizations, each with their own strengths and weaknesses. The Mercator R package facilitates selection of a biologically meaningful distance from 10 metrics, together appropriate for binary, categorical and continuous data, and visualization with 5 standard and high-dimensional graphics tools. Mercator provides a user-friendly pipeline for informaticians or biologists to perform unsupervised analyses, from exploratory pattern recognition to production of publication-quality graphics. AVAILABILITYAND IMPLEMENTATION: Mercator is freely available at the Comprehensive R Archive Network (https://cran.r-project.org/web/packages/Mercator/index.html).
Zachary B. Abrams, Caitlin E. Coombes, Suli Li, Kevin R. Coombes
Bioinform.4
2021 RCytoGPS: an R package for reading and visualizing cytogenetics data
abstract
SUMMARY: Cytogenetics data, or karyotypes, are among the most common clinically used forms of genetic data. Karyotypes are stored as standardized text strings using the International System for Human Cytogenomic Nomenclature (ISCN). Historically, these data have not been used in large-scale computational analyses due to limitations in the ISCN text format and structure. Recently developed computational tools such as CytoGPS have enabled large-scale computational analyses of karyotypes. To further enable such analyses, we have now developed RCytoGPS, an R package that takes JSON files generated from CytoGPS.org and converts them into objects in R. This conversion facilitates the analysis and visualizations of karyotype data. In effect this tool streamlines the process of performing large-scale karyotype analyses, thus advancing the field of computational cytogenetic pathology. AVAILABILITY AND IMPLEMENTATION: Freely available at https://CRAN.R-project.org/package=RCytoGPS. The code for the underlying CytoGPS software can be found at https://github.com/i2-wustl/CytoGPS.
Zachary B. Abrams, Dwayne G. Tally, Lynne V. Abruzzo, Kevin R. Coombes
Bioinform.4
2021 Pattern recognition in lymphoid malignancies using CytoGPS and Mercator
abstract
BACKGROUND: There have been many recent breakthroughs in processing and analyzing large-scale data sets in biomedical informatics. For example, the CytoGPS algorithm has enabled the use of text-based karyotypes by transforming them into a binary model. However, such advances are accompanied by new problems of data sparsity, heterogeneity, and noisiness that are magnified by the large-scale multidimensional nature of the data. To address these problems, we developed the Mercator R package, which processes and visualizes binary biomedical data. We use Mercator to address biomedical questions of cytogenetic patterns relating to lymphoid hematologic malignancies, which include a broad set of leukemias and lymphomas. Karyotype data are one of the most common form of genetic data collected on lymphoid malignancies, because karyotyping is part of the standard of care in these cancers. RESULTS: In this paper we combine the analytic power of CytoGPS and Mercator to perform a large-scale multidimensional pattern recognition study on 22,741 karyotype samples in 47 different hematologic malignancies obtained from the public Mitelman database. CONCLUSION: Our findings indicate that Mercator was able to identify both known and novel cytogenetic patterns across different lymphoid malignancies, furthering our understanding of the genetics of these diseases.
Zachary B. Abrams, Dwayne G. Tally, Lin Zhang 0056, Caitlin E. Coombes, Philip R. O. Payne, Lynne V. Abruzzo, Kevin R. Coombes
BMC Bioinform.7
2021 Simulation-derived best practices for clustering clinical data
Caitlin E. Coombes, Zachary B. Abrams, Kevin R. Coombes, Guy N. Brock
J. Biomed. Informatics4
2020 CytoGPS: A Web-Enabled Karyotype Analysis Tool for Cytogeneticists and Biomedical Data Scientists
Zachary B. Abrams, Lin Zhang 0056, Ricky Rodriguez, Lynne V. Abruzzo, Kevin R. Coombes, Philip R. P. Payne
AMIA5
2020 Unsupervised machine learning and prognostic factors of survival in chronic lymphocytic leukemia
abstract
OBJECTIVE: Unsupervised machine learning approaches hold promise for large-scale clinical data. However, the heterogeneity of clinical data raises new methodological challenges in feature selection, choosing a distance metric that captures biological meaning, and visualization. We hypothesized that clustering could discover prognostic groups from patients with chronic lymphocytic leukemia, a disease that provides biological validation through well-understood outcomes. METHODS: To address this challenge, we applied k-medoids clustering with 10 distance metrics to 2 experiments ("A" and "B") with mixed clinical features collapsed to binary vectors and visualized with both multidimensional scaling and t-stochastic neighbor embedding. To assess prognostic utility, we performed survival analysis using a Cox proportional hazard model, log-rank test, and Kaplan-Meier curves. RESULTS: In both experiments, survival analysis revealed a statistically significant association between clusters and survival outcomes (A: overall survival, P = .0164; B: time from diagnosis to treatment, P = .0039). Multidimensional scaling separated clusters along a gradient mirroring the order of overall survival. Longer survival was associated with mutated immunoglobulin heavy-chain variable region gene (IGHV) status, absent Zap 70 expression, female sex, and younger age. CONCLUSIONS: This approach to mixed-type data handling and selection of distance metric captured well-understood, binary, prognostic markers in chronic lymphocytic leukemia (sex, IGHV mutation status, ZAP70 expression status) with high fidelity.
Caitlin E. Coombes, Zachary B. Abrams, Suli Li, Lynne V. Abruzzo, Kevin R. Coombes
J. Am. Medical Informatics Assoc.5
2019 CytoGPS: a web-enabled karyotype analysis tool for cytogenetics
abstract
SUMMARY: Karyotype data are the most common form of genetic data that is regularly used clinically. They are collected as part of the standard of care in many diseases, particularly in pediatric and cancer medicine contexts. Karyotypes are represented in a unique text-based format, with a syntax defined by the International System for human Cytogenetic Nomenclature (ISCN). While human-readable, ISCN is not intrinsically machine-readable. This limitation has prevented the full use of complex karyotype data in discovery science use cases. To enhance the utility and value of karyotype data, we developed a tool named CytoGPS. CytoGPS first parses ISCN karyotypes into a machine-readable format. It then converts the ISCN karyotype into a binary Loss-Gain-Fusion (LGF) model, which represents all cytogenetic abnormalities as combinations of loss, gain, or fusion events, in a format that is analyzable using modern computational methods. Such data is then made available for comprehensive 'downstream' analyses that previously were not feasible. AVAILABILITY AND IMPLEMENTATION: Freely available at http://cytogps.org.
Zachary B. Abrams, Lin Zhang 0056, Lynne V. Abruzzo, Nyla A. Heerema, Suli Li, Tom Dillon, Ricky Rodriguez, Kevin R. Coombes, Philip R. O. Payne
Bioinform.8
2019 Inferring clonal heterogeneity in cancer using SNP arrays and whole genome sequencing
abstract
MOTIVATION: Clonal heterogeneity is common in many types of cancer, including chronic lymphocytic leukemia (CLL). Previous research suggests that the presence of multiple distinct cancer clones is associated with clinical outcome. Detection of clonal heterogeneity from high throughput data, such as sequencing or single nucleotide polymorphism (SNP) array data, is important for gaining a better understanding of cancer and may improve prediction of clinical outcome or response to treatment. Here, we present a new method, CloneSeeker, for inferring clinical heterogeneity from sequencing data, SNP array data, or both. RESULTS: We generated simulated SNP array and sequencing data and applied CloneSeeker along with two other methods. We demonstrate that CloneSeeker is more accurate than existing algorithms at determining the number of clones, distribution of cancer cells among clones, and mutation and/or copy numbers belonging to each clone. Next, we applied CloneSeeker to SNP array data from samples of 258 previously untreated CLL patients to gain a better understanding of the characteristics of CLL tumors and to elucidate the relationship between clonal heterogeneity and clinical outcome. We found that a significant majority of CLL patients appear to have multiple clones distinguished by copy number alterations alone. We also found that the presence of multiple clones corresponded with significantly worse survival among CLL patients. These findings may prove useful for improving the accuracy of prognosis and design of treatment strategies. AVAILABILITY AND IMPLEMENTATION: Code available on R-Forge: https://r-forge.r-project.org/projects/CloneSeeker/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mark R. Zucker, Lynne V. Abruzzo, Carmen D. Herling, Lynn L. Barron, Michael J. Keating, Zachary B. Abrams, Nyla A. Heerema, Kevin R. Coombes
Bioinform.8
2019 Inferring clonal heterogeneity in cancer using SNP arrays and whole genome sequencing
abstract
Bioinformatics (2019) doi: 10.1093/bioinformatics/btz057 In the original article, under heading 2.4 ‘Performance metrics’, the equations under points 2, 3 and 4 were incorrect. These have now been corrected as below.
Mark R. Zucker, Lynne V. Abruzzo, Carmen D. Herling, Lynn L. Barron, Michael J. Keating, Zachary B. Abrams, Nyla A. Heerema, Kevin R. Coombes
Bioinform.8
2019 A protocol to evaluate RNA sequencing normalization methods
abstract
BACKGROUND: RNA sequencing technologies have allowed researchers to gain a better understanding of how the transcriptome affects disease. However, sequencing technologies often unintentionally introduce experimental error into RNA sequencing data. To counteract this, normalization methods are standardly applied with the intent of reducing the non-biologically derived variability inherent in transcriptomic measurements. However, the comparative efficacy of the various normalization techniques has not been tested in a standardized manner. Here we propose tests that evaluate numerous normalization techniques and applied them to a large-scale standard data set. These tests comprise a protocol that allows researchers to measure the amount of non-biological variability which is present in any data set after normalization has been performed, a crucial step to assessing the biological validity of data following normalization. RESULTS: In this study we present two tests to assess the validity of normalization methods applied to a large-scale data set collected for systematic evaluation purposes. We tested various RNASeq normalization procedures and concluded that transcripts per million (TPM) was the best performing normalization method based on its preservation of biological signal as compared to the other methods tested. CONCLUSION: Normalization is of vital importance to accurately interpret the results of genomic and transcriptomic experiments. More work, however, needs to be performed to optimize normalization methods for RNASeq data. The present effort helps pave the way for more systematic evaluations of normalization methods across different platforms. With our proposed schema researchers can evaluate their own or future normalization methods to further improve the field of RNASeq normalization.
Zachary B. Abrams, Travis S. Johnson, Kun Huang 0001, Philip R. O. Payne, Kevin R. Coombes
BMC Bioinform.5
2018 IntLIM: integration using linear models of metabolomics and gene expression data
abstract
BACKGROUND: Integration of transcriptomic and metabolomic data improves functional interpretation of disease-related metabolomic phenotypes, and facilitates discovery of putative metabolite biomarkers and gene targets. For this reason, these data are increasingly collected in large (> 100 participants) cohorts, thereby driving a need for the development of user-friendly and open-source methods/tools for their integration. Of note, clinical/translational studies typically provide snapshot (e.g. one time point) gene and metabolite profiles and, oftentimes, most metabolites measured are not identified. Thus, in these types of studies, pathway/network approaches that take into account the complexity of transcript-metabolite relationships may neither be applicable nor readily uncover novel relationships. With this in mind, we propose a simple linear modeling approach to capture disease-(or other phenotype) specific gene-metabolite associations, with the assumption that co-regulation patterns reflect functionally related genes and metabolites. RESULTS: The proposed linear model, metabolite ~ gene + phenotype + gene:phenotype, specifically evaluates whether gene-metabolite relationships differ by phenotype, by testing whether the relationship in one phenotype is significantly different from the relationship in another phenotype (via a statistical interaction gene:phenotype p-value). Statistical interaction p-values for all possible gene-metabolite pairs are computed and significant pairs are then clustered by the directionality of associations (e.g. strong positive association in one phenotype, strong negative association in another phenotype). We implemented our approach as an R package, IntLIM, which includes a user-friendly R Shiny web interface, thereby making the integrative analyses accessible to non-computational experts. We applied IntLIM to two previously published datasets, collected in the NCI-60 cancer cell lines and in human breast tumor and non-tumor tissue, for which transcriptomic and metabolomic data are available. We demonstrate that IntLIM captures relevant tumor-specific gene-metabolite associations involved in known cancer-related pathways, including glutamine metabolism. Using IntLIM, we also uncover biologically relevant novel relationships that could be further tested experimentally. CONCLUSIONS: IntLIM provides a user-friendly, reproducible framework to integrate transcriptomic and metabolomic data and help interpret metabolomic data and uncover novel gene-metabolite relationships. The IntLIM R package is publicly available in GitHub ( https://github.com/mathelab/IntLIM ) and includes a user-friendly web application, vignettes, sample data and data/code to reproduce results.
Jalal K. Siddiqui, Elizabeth Baskin, Carmen Z. Cantemir-Stone, Bofei Zhang, Russell Bonneville, Joseph P. McElroy, Kevin R. Coombes, Ewy A. Mathé
BMC Bioinform.8
2018 Thresher: determining the number of clusters while removing outliers
abstract
BACKGROUND: Cluster analysis is the most common unsupervised method for finding hidden groups in data. Clustering presents two main challenges: (1) finding the optimal number of clusters, and (2) removing "outliers" among the objects being clustered. Few clustering algorithms currently deal directly with the outlier problem. Furthermore, existing methods for identifying the number of clusters still have some drawbacks. Thus, there is a need for a better algorithm to tackle both challenges. RESULTS: We present a new approach, implemented in an R package called Thresher, to cluster objects in general datasets. Thresher combines ideas from principal component analysis, outlier filtering, and von Mises-Fisher mixture models in order to select the optimal number of clusters. We performed a large Monte Carlo simulation study to compare Thresher with other methods for detecting outliers and determining the number of clusters. We found that Thresher had good sensitivity and specificity for detecting and removing outliers. We also found that Thresher is the best method for estimating the optimal number of clusters when the number of objects being clustered is smaller than the number of variables used for clustering. Finally, we applied Thresher and eleven other methods to 25 sets of breast cancer data downloaded from the Gene Expression Omnibus; only Thresher consistently estimated the number of clusters to lie in the range of 4-7 that is consistent with the literature. CONCLUSIONS: Thresher is effective at automatically detecting and removing outliers. By thus cleaning the data, it produces better estimates of the optimal number of clusters when there are more variables than objects. When we applied Thresher to a variety of breast cancer datasets, it produced estimates that were both self-consistent and consistent with the literature. We expect Thresher to be useful for studying a wide variety of biological datasets.
Zachary B. Abrams, Steven M. Kornblau, Kevin R. Coombes
BMC Bioinform.4
2015 Development of a robust classifier for quality control of reverse-phase protein arrays
abstract
MOTIVATION: High-throughput reverse-phase protein array (RPPA) technology allows for the parallel measurement of protein expression levels in approximately 1000 samples. However, the many steps required in the complex protocol (sample lysate preparation, slide printing, hybridization, washing and amplified detection) may create substantial variability in data quality. We are not aware of any other quality control algorithm that is tuned to the special characteristics of RPPAs. RESULTS: We have developed a novel classifier for quality control of RPPA experiments using a generalized linear model and logistic function. The outcome of the classifier, ranging from 0 to 1, is defined as the probability that a slide is of good quality. After training, we tested the classifier using two independent validation datasets. We conclude that the classifier can distinguish RPPA slides of good quality from those of poor quality sufficiently well such that normalization schemes, protein expression patterns and advanced biological analyses will not be drastically impacted by erroneous measurements or systematic variations. AVAILABILITY AND IMPLEMENTATION: The classifier, implemented in the "SuperCurve" R package, can be freely downloaded at http://bioinformatics.mdanderson.org/main/OOMPA:Overview or http://r-forge.r-project.org/projects/supercurve/. The data used to develop and validate the classifier are available at http://bioinformatics.mdanderson.org/MOAR.
Zhenlin Ju, Paul L. Roebuck, Doris R. Siwak, Nianxiang Zhang, Yiling Lu, Michael A. Davies, Rehan Akbani, John N. Weinstein, Gordon B. Mills, Kevin R. Coombes
Bioinform.11
2015 drexplorer: A tool to explore dose-response relationships and drug-drug interactions
abstract
Abstract Motivation: Nonlinear dose–response models are primary tools for estimating the potency [e.g. half-maximum inhibitory concentration (IC) known as IC50] of anti-cancer drugs. We present drexplorer software, which enables biologists to evaluate replicate reproducibility, detect outlier data points, fit different models, select the best model, estimate IC values at different percentiles and assess drug–drug interactions. drexplorer serves as a computation engine within the R environment and a graphical interface for users who do not have programming backgrounds. Availability and implementation: The drexplorer R package is freely available from GitHub at https://github.com/nickytong/drexplorer. A graphical user interface is shipped with the package. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Pan Tong, Kevin R. Coombes, Faye M. Johnson, Lauren A. Byers, Lixia Diao, Diane D. Liu, J. Jack Lee, John V. Heymach, Jing Wang 0006
Bioinform.2
2014 Latent Feature Decompositions for Integrative Analysis of Multi-Platform Genomic Data
abstract
Increased availability of multi-platform genomics data on matched samples has sparked research efforts to discover how diverse molecular features interact both within and between platforms. In addition, simultaneous measurements of genetic and epigenetic characteristics illuminate the roles their complex relationships play in disease progression and outcomes. However, integrative methods for diverse genomics data are faced with the challenges of ultra-high dimensionality and the existence of complex interactions both within and between platforms. We propose a novel modeling framework for integrative analysis based on decompositions of the large number of platform-specific features into a smaller number of latent features. Subsequently we build a predictive model for clinical outcomes accounting for both within- and between-platform interactions based on Bayesian model averaging procedures. Principal components, partial least squares and non-negative matrix factorization as well as sparse counterparts of each are used to define the latent features, and the performance of these decompositions is compared both on real and simulated data. The latent feature interactions are shown to preserve interactions between the original features and not only aid prediction but also allow explicit selection of outcome-related features. The methods are motivated by and applied to a glioblastoma multiforme data set from The Cancer Genome Atlas to predict patient survival times integrating gene expression, microRNA, copy number and methylation data. For the glioblastoma data, we find a high concordance between our selected prognostic genes and genes with known associations with glioblastoma. In addition, our model discovers several relevant cross-platform interactions such as copy number variation associated gene dosing and epigenetic regulation through promoter methylation. On simulated data, we show that our proposed method successfully incorporates interactions within and between genomic platforms to aid accurate prediction and variable selection. Our methods perform best when principal components are used to define the latent features.
Karl B. Gregory, Amin A. Momin, Kevin R. Coombes, Veerabhadran Baladandayuthapani
IEEE ACM Trans. Comput. Biol. Bioinform.3
2013 targetHub: a programmable interface for miRNA-gene interactions
abstract
MOTIVATION: With the expansion of high-throughput technologies, understanding different kinds of genome-level data is a common task. MicroRNA (miRNA) is increasingly profiled using high-throughput technologies (microarrays or next-generation sequencing). The downstream analysis of miRNA targets can be difficult. Although there are many databases and algorithms to predict miRNA targets, there are few tools to integrate miRNA-gene interaction data into high-throughput genomic analyses. RESULTS: We present targetHub, a CouchDB database of miRNA-gene interactions. TargetHub provides a programmer-friendly interface to access miRNA targets. The Web site provides RESTful access to miRNA-gene interactions with an assortment of gene and miRNA identifiers. It can be a useful tool to integrate miRNA target interaction data directly into high-throughput bioinformatics analyses. AVAILABILITY: TargetHub is available on the web at http://app1.bioinformatics.mdanderson.org/tarhub/_design/basic/index.html.
Ganiraju Manyam, Cristina Ivan, George A. Calin, Kevin R. Coombes
Bioinform.4
2013 SIBER: systematic identification of bimodally expressed genes using RNAseq data
abstract
MOTIVATION: Identification of bimodally expressed genes is an important task, as genes with bimodal expression play important roles in cell differentiation, signalling and disease progression. Several useful algorithms have been developed to identify bimodal genes from microarray data. Currently, no method can deal with data from next-generation sequencing, which is emerging as a replacement technology for microarrays. RESULTS: We present SIBER (systematic identification of bimodally expressed genes using RNAseq data) for effectively identifying bimodally expressed genes from next-generation RNAseq data. We evaluate several candidate methods for modelling RNAseq count data and compare their performance in identifying bimodal genes through both simulation and real data analysis. We show that the lognormal mixture model performs best in terms of power and robustness under various scenarios. We also compare our method with alternative approaches, including profile analysis using clustering and kurtosis (PACK) and cancer outlier profile analysis (COPA). Our method is robust, powerful, invariant to shifting and scaling, has no blind spots and has a sample-size-free interpretation. AVAILABILITY: The R package SIBER is available at the website http://bioinformatics.mdanderson.org/main/OOMPA:Overview.
Pan Tong, Yong Chen 0016, Kevin R. Coombes
Bioinform.4
2012 integIRTy: a method to identify genes altered in cancer by accounting for multiple mechanisms of regulation using item response theory
abstract
MOTIVATION: Identifying genes altered in cancer plays a crucial role in both understanding the mechanism of carcinogenesis and developing novel therapeutics. It is known that there are various mechanisms of regulation that can lead to gene dysfunction, including copy number change, methylation, abnormal expression, mutation and so on. Nowadays, all these types of alterations can be simultaneously interrogated by different types of assays. Although many methods have been proposed to identify altered genes from a single assay, there is no method that can deal with multiple assays accounting for different alteration types systematically. RESULTS: In this article, we propose a novel method, integration using item response theory (integIRTy), to identify altered genes by using item response theory that allows integrated analysis of multiple high-throughput assays. When applied to a single assay, the proposed method is more robust and reliable than conventional methods such as Student's t-test or the Wilcoxon rank-sum test. When used to integrate multiple assays, integIRTy can identify novel-altered genes that cannot be found by looking at individual assay separately. We applied integIRTy to three public cancer datasets (ovarian carcinoma, breast cancer, glioblastoma) for cross-assay type integration which all show encouraging results. AVAILABILITY AND IMPLEMENTATION: The R package integIRTy is available at the web site http://bioinformatics.mdanderson.org/main/OOMPA:Overview. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pan Tong, Kevin R. Coombes
Bioinform.2
2012 Sources of variation in false discovery rate estimation include sample size, correlation, and inherent differences between groups
abstract
BACKGROUND: High-throughtput technologies enable the testing of tens of thousands of measurements simultaneously. Identification of genes that are differentially expressed or associated with clinical outcomes invokes the multiple testing problem. False Discovery Rate (FDR) control is a statistical method used to correct for multiple comparisons for independent or weakly dependent test statistics. Although FDR control is frequently applied to microarray data analysis, gene expression is usually correlated, which might lead to inaccurate estimates. In this paper, we evaluate the accuracy of FDR estimation. METHODS: Using two real data sets, we resampled subgroups of patients and recalculated statistics of interest to illustrate the imprecision of FDR estimation. Next, we generated many simulated data sets with block correlation structures and realistic noise parameters, using the Ultimate Microarray Prediction, Inference, and Reality Engine (UMPIRE) R package. We estimated FDR using a beta-uniform mixture (BUM) model, and examined the variation in FDR estimation. RESULTS: The three major sources of variation in FDR estimation are the sample size, correlations among genes, and the true proportion of differentially expressed genes (DEGs). The sample size and proportion of DEGs affect both magnitude and precision of FDR estimation, while the correlation structure mainly affects the variation of the estimated parameters. CONCLUSIONS: We have decomposed various factors that affect FDR estimation, and illustrated the direction and extent of the impact. We found that the proportion of DEGs has a significant impact on FDR; this factor might have been overlooked in previous studies and deserves more thought when controlling FDR.
Jiexin Zhang 0004, Kevin R. Coombes
BMC Bioinform.2
2009 Variable slope normalization of reverse phase protein arrays
abstract
MOTIVATION: Reverse phase protein arrays (RPPA) measure the relative expression levels of a protein in many samples simultaneously. A set of identically spotted arrays can be used to measure the levels of more than one protein. Protein expression within each sample on an array is estimated by borrowing strength across all the samples, but using only within array information. When comparing across slides, it is essential to account for sample loading, the total amount of protein printed per sample. Currently, total protein is estimated using either a housekeeping protein or the sample median across all slides. When the variability in sample loading is large, these methods are suboptimal because they do not account for the fact that the protein expression for each slide is estimated separately. RESULTS: We propose a new normalization method for RPPA data, called variable slope (VS) normalization, that takes into account that quantification of RPPA slides is performed separately. This method is better able to remove loading bias and recover true correlation structures between proteins. AVAILABILITY: Code to implement the method in the statistical package R and anonymized data are available at (http://bioinformatics.mdanderson.org/supplements.html).
E. Shannon Neeley, Steven M. Kornblau, Kevin R. Coombes, Keith A. Baggerly
Bioinform.3
2009 Serial dilution curve: a new method for analysis of reverse phase protein array data
abstract
UNLABELLED: Reverse phase protein arrays (RPPAs) are a powerful high-throughput tool for measuring protein concentrations in a large number of samples. In RPPA technology, the original samples are often diluted successively multiple times, forming dilution series to extend the dynamic range of the measurements and to increase confidence in quantitation. An RPPA experiment is equivalent to running multiple ELISA assays concurrently except that there is usually no known protein concentration from which one can construct a standard response curve. Here, we describe a new method called 'serial dilution curve for RPPA data analysis'. Compared with the existing methods, the new method has the advantage of using fewer parameters and offering a simple way of visualizing the raw data. We showed how the method can be used to examine data quality and to obtain robust quantification of protein concentrations. AVAILABILITY: A computer program in R for using serial dilution curve for RPPA data analysis is freely available at http://odin.mdacc.tmc.edu/~zhangli/RPPA.
Qingyi Wei, Gordon B. Mills, Kevin R. Coombes
Bioinform.6
2008 RPPAML/RIMS: A metadata format and an information management system for reverse phase protein arrays
abstract
BACKGROUND: Reverse Phase Protein Arrays (RPPA) are convenient assay platforms to investigate the presence of biomarkers in tissue lysates. As with other high-throughput technologies, substantial amounts of analytical data are generated. Over 1,000 samples may be printed on a single nitrocellulose slide. Up to 100 different proteins may be assessed using immunoperoxidase or immunoflorescence techniques in order to determine relative amounts of protein expression in the samples of interest. RESULTS: In this report an RPPA Information Management System (RIMS) is described and made available with open source software. In order to implement the proposed system, we propose a metadata format known as reverse phase protein array markup language (RPPAML). RPPAML would enable researchers to describe, document and disseminate RPPA data. The complexity of the data structure needed to describe the results and the graphic tools necessary to visualize them require a software deployment distributed between a client and a server application. This was achieved without sacrificing interoperability between individual deployments through the use of an open source semantic database, S3DB. This data service backbone is available to multiple client side applications that can also access other server side deployments. The RIMS platform was designed to interoperate with other data analysis and data visualization tools such as Cytoscape. CONCLUSION: The proposed RPPAML data format hopes to standardize RPPA data. Standardization of data would result in diverse client applications being able to operate on the same set of data. Additionally, having data in a standard format would enable data dissemination and data analysis.
Romesh Stanislaus, Mark Carey, Helena F. Deus, Kevin R. Coombes, Bryan T. J. Hennessy, Gordon B. Mills, Jonas S. Almeida
BMC Bioinform.4
2007 Enrichment analysis in high-throughput genomics - accounting for dependency in the NULL
abstract
Translating the overwhelming amount of data generated in high-throughput genomics experiments into biologically meaningful evidence, which may for example point to a series of biomarkers or hint at a relevant pathway, is a matter of great interest in bioinformatics these days. Genes showing similar experimental profiles, it is hypothesized, share biological mechanisms that if understood could provide clues to the molecular processes leading to pathological events. It is the topic of further study to learn if or how a priori information about the known genes may serve to explain coexpression. One popular method of knowledge discovery in high-throughput genomics experiments, enrichment analysis (EA), seeks to infer if an interesting collection of genes is 'enriched' for a Consortium particular set of a priori Gene Ontology Consortium (GO) classes. For the purposes of statistical testing, the conventional methods offered in EA software implicitly assume independence between the GO classes. Genes may be annotated for more than one biological classification, and therefore the resulting test statistics of enrichment between GO classes can be highly dependent if the overlapping gene sets are relatively large. There is a need to formally determine if conventional EA results are robust to the independence assumption. We derive the exact null distribution for testing enrichment of GO classes by relaxing the independence assumption using well-known statistical theory. In applications with publicly available data sets, our test results are similar to the conventional approach which assumes independence. We argue that the independence assumption is not detrimental.
David L. Gold, Kevin R. Coombes, Jing Wang 0006, Bani K. Mallick
Briefings Bioinform.2
2007 Non-parametric quantification of protein lysate arrays
abstract
MOTIVATION: Proteins play a crucial role in biological activity, so much can be learned from measuring protein expression and post-translational modification quantitatively. The reverse-phase protein lysate arrays allow us to quantify the relative expression levels of a protein in many different cellular samples simultaneously. Existing approaches to quantify protein arrays use parametric response curves fit to dilution series data. The results can be biased when the parametric function does not fit the data. RESULTS: We propose a non-parametric approach which adapts to any monotone response curve. The non-parametric approach is shown to be promising via both simulation and real data studies; it reduces the bias due to model misspecification and protects against outliers in the data. The non-parametric approach enables more reliable quantification of protein lysate arrays. AVAILABILITY: Code to implement the proposed method in the statistical package R is available at: http://odin.mdacc.tmc.edu/jhu/lysatearray-analysis/
Xuming He 0002, Keith A. Baggerly, Kevin R. Coombes, Bryan T. J. Hennessy, Gordon B. Mills
Bioinform.4
2007 PrepMS: TOF MS data graphical preprocessing tool
abstract
UNLABELLED: We introduce a simple-to-use graphical tool that enables researchers to easily prepare time-of-flight mass spectrometry data for analysis. For ease of use, the graphical executable provides default parameter settings, experimentally determined to work well in most situations. These values, if desired, can be changed by the user. PrepMS is a stand-alone application made freely available (open source), and is under the General Public License (GPL). Its graphical user interface, default parameter settings, and display plots allow PrepMS to be used effectively for data preprocessing, peak detection and visual data quality assessment. AVAILABILITY: Stand-alone executable files and Matlab toolbox are available for download at: http://sourceforge.net/projects/prepms
Yuliya V. Karpievitch, Elizabeth G. Hill, Adam J. Smolka, Jeffrey S. Morris, Kevin R. Coombes, Keith A. Baggerly, Jonas S. Almeida
Bioinform.5
2006 Gene sequence signatures revealed by mining the UniGene affiliation network
abstract
BACKGROUND: In the post-genomic era, developing tools to decode biological information from genomic sequences is important. Inspired by affiliation network theory, we investigated gene sequences of two kinds of UniGene clusters (UCs): narrowly expressed transcripts (NETs), whose expression is confined to a few tissues; and prevalently expressed transcripts (PETs) that are expressed in many tissues. RESULTS: We explored the human and the mouse UniGene databases to compare NETs and PETs from different perspectives. We found that NETs were associated with smaller cluster size, shorter sequence length, a lower likelihood of having LocusLink annotations, and lower and more sporadic levels of expression. Significantly, the dinucleotide frequencies of NETs are similar to those of intergenic sequences in the genome, and they differ from those of PETs. We used these differences in dinucleotide frequencies to develop a discriminant analysis model to distinguish PETs from intergenic sequences. CONCLUSIONS: Our results show that most NETs resemble intergenic sequences, casting doubts on the quality of such UniGene clusters. However, we also noted that a fraction of NETs resemble PETs in terms of dinucleotide frequencies and other features. Such NETs may have fewer quality problems. This work may be helpful in the studies of non-coding RNAs and in the validation of gene sequence databases.
Jiexin Zhang 0004, Kevin R. Coombes
Bioinform.3
2005 Analysis of dose-response effects on gene expression data with comparison of two microarray platforms
abstract
MOTIVATION: The problems of analyzing dose effects on gene expression are gaining attention in biomedical research. A specific challenge is to detect genes with expression levels that change according to dose levels in a non-random manner, but nonetheless may be considered as potential biomarkers. METHOD: We are among the first to formally apply a tool that uses an isotonic (monotonic) regression approach to this area of study. We introduce a test statistic to select genes with significant dose-response expression in a monotonic fashion based on a permutation procedure. We then compare the results with those achieved from the application of a likelihood ratio-based test. RESULTS: We apply the isotonic regression approach to a study of gene expression in the RKO colon carcinoma cell line in response to varying dosage levels of the chemotherapeutic agent 5-fluorouracil. A feature of both Affymetrix and printed 75mer oligomer cDNA arrays produced from the same samples provides an opportunity to compare the two microarray platforms. AVAILABILITY: Statistical software S-plus Code to implement the method is available from the authors. CONTACT: [email protected]
Mini Kapoor, Wei Zhang 0011, Stanley R. Hamilton, Kevin R. Coombes
Bioinform.5
2005 Applications of beta-mixture models in bioinformatics
abstract
SUMMARY: We propose a beta-mixture model approach to solve a variety of problems related to correlations of gene-expression levels. For example, in meta-analyses of microarray gene-expression datasets, a threshold value of correlation coefficients for gene-expression levels is used to decide whether gene-expression levels are strongly correlated across studies. Ad hoc threshold values such as 0.5 are often used. In this paper, we use a beta-mixture model approach to divide the correlation coefficients into several populations so that the large correlation coefficients can be identified. Another important application of the proposed method is in finding co-expressed genes. Two examples are provided to illustrate both applications. Through our analysis, we also discover that the popular model selection criteria BIC and AIC are not suitable for the beta-mixture model. To determine the number of components in the mixture model, we suggest an alternative criterion, ICL-BIC, which is shown to perform better in selecting the correct mixture model. SUPPLEMENTARY INFORMATION: http://odin.mdacc.tmc.edu/~yuanj/highcorgeneanno.html.
Chunlei Wu, Jing Wang 0006, Kevin R. Coombes
Bioinform.5
2005 Feature extraction and quantification for mass spectrometry in biomedical applications using the mean spectrum
abstract
MOTIVATION: Mass spectrometry yields complex functional data for which the features of scientific interest are peaks. A common two-step approach to analyzing these data involves first extracting and quantifying the peaks, then analyzing the resulting matrix of peak quantifications. Feature extraction and quantification involves a number of interrelated steps. It is important to perform these steps well, since subsequent analyses condition on these determinations. Also, it is difficult to compare the performance of competing methods for analyzing mass spectrometry data since the true expression levels of the proteins in the population are generally not known. RESULTS: In this paper, we introduce a new method for feature extraction in mass spectrometry data that uses translation-invariant wavelet transforms and performs peak detection using the mean spectrum. We examine the method's performance through examples and simulation, and demonstrate the advantages of using the mean spectrum to detect peaks. We also describe a new physics-based computer model of mass spectrometry and demonstrate how one may design simulation studies based on this tool to systematically compare competing methods. AVAILABILITY: MATLAB scripts to implement the methods described in this paper and R code for the virtual mass spectrometer are available at http://bioinformatics.mdanderson.org/software.html SUPPLEMENTARY INFORMATION: http://bioinformatics.mdanderson.org/supplements.html.
Jeffrey S. Morris, Kevin R. Coombes, John M. Koomen, Keith A. Baggerly, Ryuji Kobayashi
Bioinform.2
2004 Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments
abstract
MOTIVATION: There has been much interest in using patterns derived from surface-enhanced laser desorption and ionization (SELDI) protein mass spectra from serum to differentiate samples from patients both with and without disease. Such patterns have been used without identification of the underlying proteins responsible. However, there are questions as to the stability of this procedure over multiple experiments. RESULTS: We compared SELDI proteomic spectra from serum from three experiments by the same group on separating ovarian cancer from normal tissue. These spectra are available on the web at http://clinicalproteomics.steem.com. In general, the results were not reproducible across experiments. Baseline correction prevents reproduction of the results for two of the experiments. In one experiment, there is evidence of a major shift in protocol mid-experiment which could bias the results. In another, structure in the noise regions of the spectra allows us to distinguish normal from cancer, suggesting that the normals and cancers were processed differently. Sets of features found to discriminate well in one experiment do not generalize to other experiments. Finally, the mass calibration in all three experiments appears suspect. Taken together, these and other concerns suggest that much of the structure uncovered in these experiments could be due to artifacts of sample processing, not to the underlying biology of cancer. We provide some guidelines for design and analysis in experiments like these to ensure better reproducible, biologically meaningfully results. AVAILABILITY: The MATLAB and Perl code used in our analyses is available at http://bioinformatics.mdanderson.org
Keith A. Baggerly, Jeffrey S. Morris, Kevin R. Coombes
Bioinform.3
2004 Differences in gene expression between B-cell chronic lymphocytic leukemia and normal B cells: a meta-analysis of three microarray studies
abstract
Abstract Motivation: A major focus of current cancer research is to identify genes that can be used as markers for prognosis and diagnosis, and as targets for therapy. Microarray technology has been applied extensively for this purpose, even though it has been reported that the agreement between microarray platforms is poor. A critical question is: how can we best combine the measurements of matched genes across microarray platforms to develop diagnostic and prognostic tools related to the underlying biology? Results: We introduce a statistical approach within a Bayesian framework to combine the microarray data on matched genes from three investigations of gene expression profiling of B-cell chronic lymphocytic leukemia (CLL) and normal B cells (NBC) using three different microarray platforms, oligonucleotide arrays, cDNA arrays printed on glass slides and cDNA arrays printed on nylon membranes. Using this approach, we identified a number of genes that were consistently differentially expressed between CLL and NBC samples. Availability: Glass slide cDNA array data are available through the public archives of Stanford University at http://cmgm.stanford.edu/pbrown/. Oligonucleotide array data are available in the supplemental material from Klein ηl. Nylon membrane cDNA microarray data are freely available at our website, http://bioinformatics.mdanderson.org/pubdata.html. Software to batch-process gene annotations (GeneLink) is available at http://bioinformatics.mdanderson.org/GeneLink.html. Image quantification software (ScanAlyze) is available at http://rana.lbl.gov/downloads/ScanAlyze.zip. Statistical software (S-Plus 2000) is available commercially from Insightful Corp., Seattle, WA. Software to determine the co-occurrence of terms in PubMed abstracts (PDQ_MED) is commercially available from Inpharmix Inc., Greenwood, IN.
Jing Wang 0006, Kevin R. Coombes, William Edward Highsmith, Michael J. Keating, Lynne V. Abruzzo
Bioinform.2