Keith A. Baggerly

dblp:88/5938 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 97% Computational science and engineering · 3%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
epigenomics
0.512021
Methylation-eQTL analysis in cancer research · Bioinform. 2021
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.232007
PrepMS: TOF MS data graphical preprocessing tool · Bioinform. 2007
Feature extraction and quantification for mass spectrometry in biomedical applications using the mean spectrum · Bioinform. 2005
Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments · Bioinform. 2004
Bioinformatics and computational biology
proteomics
0.232007
PrepMS: TOF MS data graphical preprocessing tool · Bioinform. 2007
Feature extraction and quantification for mass spectrometry in biomedical applications using the mean spectrum · Bioinform. 2005
Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments · Bioinform. 2004
Bioinformatics and computational biology › epigenomics › ChIP-seq analysis
peak detection
0.122007
Feature extraction and quantification for mass spectrometry in biomedical applications using the mean spectrum · Bioinform. 2005
PrepMS: TOF MS data graphical preprocessing tool · Bioinform. 2007
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.112007
Short oligonucleotide probes containing G-stacks display abnormal binding affinity on Affymetrix microarrays · Bioinform. 2007
Bioinformatics and computational biology
microarray probe design
0.112007
Short oligonucleotide probes containing G-stacks display abnormal binding affinity on Affymetrix microarrays · Bioinform. 2007
Bioinformatics and computational biology › proteomics
protein quantification
0.112007
Non-parametric quantification of protein lysate arrays · Bioinform. 2007
Bioinformatics and computational biology › proteomics › mass spectrometry data analysis
spectral preprocessing
0.112007
PrepMS: TOF MS data graphical preprocessing tool · Bioinform. 2007
Computational science and engineering › computational reproducibility
reproducibility assessment
0.012004
Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments · Bioinform. 2004
Bioinformatics and computational biology › gene expression analysis
differential expression analysis
0.012003
Differential expression in SAGE: accounting for normal between-library variation · Bioinform. 2003
Bioinformatics and computational biology
gene expression analysis
0.012003
Differential expression in SAGE: accounting for normal between-library variation · Bioinform. 2003
Bioinformatics and computational biology › proteomics
reverse-phase protein array analysis
0.012007
Non-parametric quantification of protein lysate arrays · Bioinform. 2007
Bioinformatics and computational biology
biomarker discovery
0.012004
Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments · Bioinform. 2004
Bioinformatics and computational biology
overdispersion modeling
0.012003
Differential expression in SAGE: accounting for normal between-library variation · Bioinform. 2003
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.012003
Differential expression in SAGE: accounting for normal between-library variation · Bioinform. 2003

Methods — techniques the papers use, named apart from their topics

sequential regression · 0.5penalized regression · 0.5variable slope normalization · 0.1loading bias correction · 0.1time-of-flight mass spectrometry · 0.1positional dependent nearest neighbor model · 0.1nonparametric regression · 0.1isotonic regression · 0.1simulation modeling · 0.1mean spectrum · 0.1
YearPublicationVenuePosition
2021 Methylation-eQTL analysis in cancer research
abstract
MOTIVATION: DNA methylation is a key epigenetic factor regulating gene expression. While promoter methylation has been well studied, recent publications have revealed that functionally important methylation also occurs in intergenic and distal regions, and varies across genes and tissue types. Given the growing importance of inter-platform integrative genomic analyses, there is an urgent need to develop methods to discover and characterize gene-level relationships between methylation and expression. RESULTS: We introduce a novel sequential penalized regression approach to identify methylation-expression quantitative trait loci (methyl-eQTLs), a term that we have coined to represent, for each gene and tissue type, a sparse set of CpG loci best explaining gene expression and accompanying weights indicating direction and strength of association. Using TCGA and MD Anderson colorectal cohorts to build and validate our models, we demonstrate our strategy better explains expression variability than current commonly used gene-level methylation summaries. The methyl-eQTLs identified by our approach can be used to construct gene-level methylation summaries that are maximally correlated with gene expression for use in integrative models, and produce a tissue-specific summary of which genes appear to be strongly regulated by methylation. Our results introduce an important resource to the biomedical community for integrative genomics analyses involving DNA methylation. AVAILABILITY AND IMPLEMENTATION: We produce an R Shiny app (https://rstudio-prd-c1.pmacs.upenn.edu/methyl-eQTL/) that interactively presents methyl-eQTL results for colorectal, breast and pancreatic cancer. The source R code for this work is provided in the Supplementary Material. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yusha Liu, Keith A. Baggerly, Elias Orouji, Ganiraju Manyam, Michael Lam, Jennifer S. Davis, Michael S. Lee, Bradley M. Broom, David G. Menter, Kunal Rai, Scott Kopetz, Jeffrey S. Morris
Bioinform.2
2009 Variable slope normalization of reverse phase protein arrays
abstract
MOTIVATION: Reverse phase protein arrays (RPPA) measure the relative expression levels of a protein in many samples simultaneously. A set of identically spotted arrays can be used to measure the levels of more than one protein. Protein expression within each sample on an array is estimated by borrowing strength across all the samples, but using only within array information. When comparing across slides, it is essential to account for sample loading, the total amount of protein printed per sample. Currently, total protein is estimated using either a housekeeping protein or the sample median across all slides. When the variability in sample loading is large, these methods are suboptimal because they do not account for the fact that the protein expression for each slide is estimated separately. RESULTS: We propose a new normalization method for RPPA data, called variable slope (VS) normalization, that takes into account that quantification of RPPA slides is performed separately. This method is better able to remove loading bias and recover true correlation structures between proteins. AVAILABILITY: Code to implement the method in the statistical package R and anonymized data are available at (http://bioinformatics.mdanderson.org/supplements.html).
E. Shannon Neeley, Steven M. Kornblau, Kevin R. Coombes, Keith A. Baggerly
Bioinform.4
2007 Non-parametric quantification of protein lysate arrays
abstract
MOTIVATION: Proteins play a crucial role in biological activity, so much can be learned from measuring protein expression and post-translational modification quantitatively. The reverse-phase protein lysate arrays allow us to quantify the relative expression levels of a protein in many different cellular samples simultaneously. Existing approaches to quantify protein arrays use parametric response curves fit to dilution series data. The results can be biased when the parametric function does not fit the data. RESULTS: We propose a non-parametric approach which adapts to any monotone response curve. The non-parametric approach is shown to be promising via both simulation and real data studies; it reduces the bias due to model misspecification and protects against outliers in the data. The non-parametric approach enables more reliable quantification of protein lysate arrays. AVAILABILITY: Code to implement the proposed method in the statistical package R is available at: http://odin.mdacc.tmc.edu/jhu/lysatearray-analysis/
Xuming He 0002, Keith A. Baggerly, Kevin R. Coombes, Bryan T. J. Hennessy, Gordon B. Mills
Bioinform.3
2007 PrepMS: TOF MS data graphical preprocessing tool
abstract
UNLABELLED: We introduce a simple-to-use graphical tool that enables researchers to easily prepare time-of-flight mass spectrometry data for analysis. For ease of use, the graphical executable provides default parameter settings, experimentally determined to work well in most situations. These values, if desired, can be changed by the user. PrepMS is a stand-alone application made freely available (open source), and is under the General Public License (GPL). Its graphical user interface, default parameter settings, and display plots allow PrepMS to be used effectively for data preprocessing, peak detection and visual data quality assessment. AVAILABILITY: Stand-alone executable files and Matlab toolbox are available for download at: http://sourceforge.net/projects/prepms
Yuliya V. Karpievitch, Elizabeth G. Hill, Adam J. Smolka, Jeffrey S. Morris, Kevin R. Coombes, Keith A. Baggerly, Jonas S. Almeida
Bioinform.6
2007 Short oligonucleotide probes containing G-stacks display abnormal binding affinity on Affymetrix microarrays
abstract
MOTIVATION: In microarray experiments, probe design is critical to the specific and accurate measurement of target concentrations. Current designs select suitable probes through in silico scanning of transcriptome/genome based on first principles. However, due to lack of tools, the observed microarray data have not been used to assess the performance of individual probes to provide feedback to improve future designs. RESULT: In this study, we describe a probe performance assessment method based on the concordance of the observed signals from probes that share common targets. Using this method, we found that probes containing multiple guanines in a row (G-stacks) have abnormal binding behavior compared with other probes, both in gene expression assays and genotyping assays using Affymetrix microarrays. These probes are less likely to covary with other probes that interrogate the same genes. Moreover, we found that these probes are much more likely to produce outliers when fitting the observed signals according to the positional dependent nearest neighbor model, which gives reasonable estimates of binding affinity for most other probes. These results suggest that probes containing G-stacks tend to have increased cross hybridization signals and reduced target-specific hybridization signals, presumably due to multiplex binding forming G-quartet structures. Our findings are expected to be useful in microarray design and data analysis.
Chunlei Wu, Keith A. Baggerly, Roberto Carta
Bioinform.3
2005 Feature extraction and quantification for mass spectrometry in biomedical applications using the mean spectrum
abstract
MOTIVATION: Mass spectrometry yields complex functional data for which the features of scientific interest are peaks. A common two-step approach to analyzing these data involves first extracting and quantifying the peaks, then analyzing the resulting matrix of peak quantifications. Feature extraction and quantification involves a number of interrelated steps. It is important to perform these steps well, since subsequent analyses condition on these determinations. Also, it is difficult to compare the performance of competing methods for analyzing mass spectrometry data since the true expression levels of the proteins in the population are generally not known. RESULTS: In this paper, we introduce a new method for feature extraction in mass spectrometry data that uses translation-invariant wavelet transforms and performs peak detection using the mean spectrum. We examine the method's performance through examples and simulation, and demonstrate the advantages of using the mean spectrum to detect peaks. We also describe a new physics-based computer model of mass spectrometry and demonstrate how one may design simulation studies based on this tool to systematically compare competing methods. AVAILABILITY: MATLAB scripts to implement the methods described in this paper and R code for the virtual mass spectrometer are available at http://bioinformatics.mdanderson.org/software.html SUPPLEMENTARY INFORMATION: http://bioinformatics.mdanderson.org/supplements.html.
Jeffrey S. Morris, Kevin R. Coombes, John M. Koomen, Keith A. Baggerly, Ryuji Kobayashi
Bioinform.4
2004 Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments
abstract
MOTIVATION: There has been much interest in using patterns derived from surface-enhanced laser desorption and ionization (SELDI) protein mass spectra from serum to differentiate samples from patients both with and without disease. Such patterns have been used without identification of the underlying proteins responsible. However, there are questions as to the stability of this procedure over multiple experiments. RESULTS: We compared SELDI proteomic spectra from serum from three experiments by the same group on separating ovarian cancer from normal tissue. These spectra are available on the web at http://clinicalproteomics.steem.com. In general, the results were not reproducible across experiments. Baseline correction prevents reproduction of the results for two of the experiments. In one experiment, there is evidence of a major shift in protocol mid-experiment which could bias the results. In another, structure in the noise regions of the spectra allows us to distinguish normal from cancer, suggesting that the normals and cancers were processed differently. Sets of features found to discriminate well in one experiment do not generalize to other experiments. Finally, the mass calibration in all three experiments appears suspect. Taken together, these and other concerns suggest that much of the structure uncovered in these experiments could be due to artifacts of sample processing, not to the underlying biology of cancer. We provide some guidelines for design and analysis in experiments like these to ensure better reproducible, biologically meaningfully results. AVAILABILITY: The MATLAB and Perl code used in our analyses is available at http://bioinformatics.mdanderson.org
Keith A. Baggerly, Jeffrey S. Morris, Kevin R. Coombes
Bioinform.1
2004 Overdispersed logistic regression for SAGE: Modelling multiple groups and covariates
abstract
BACKGROUND: Two major identifiable sources of variation in data derived from the Serial Analysis of Gene Expression (SAGE) are within-library sampling variability and between-library heterogeneity within a group. Most published methods for identifying differential expression focus on just the sampling variability. In recent work, the problem of assessing differential expression between two groups of SAGE libraries has been addressed by introducing a beta-binomial hierarchical model that explicitly deals with both of the above sources of variation. This model leads to a test statistic analogous to a weighted two-sample t-test. When the number of groups involved is more than two, however, a more general approach is needed. RESULTS: We describe how logistic regression with overdispersion supplies this generalization, carrying with it the framework for incorporating other covariates into the model as a byproduct. This approach has the advantage that logistic regression routines are available in several common statistical packages. CONCLUSIONS: The described method provides an easily implemented tool for analyzing SAGE data that correctly handles multiple types of variation and allows for more flexible modelling.
Keith A. Baggerly, Jeffrey S. Morris, C. Marcelo Aldaz
BMC Bioinform.1
2003 Differential expression in SAGE: accounting for normal between-library variation
abstract
Abstract Motivation: In contrasting levels of gene expression between groups of SAGE libraries, the libraries within each group are often combined and the counts for the tag of interest summed, and inference is made on the basis of these larger ‘pseudolibraries’. While this captures the sampling variability inherent in the procedure, it fails to allow for normal variation in levels of the gene between individuals within the same group, and can consequently overstate the significance of the results. The effect is not slight: between-library variation can be hundreds of times the within-library variation. Results: We introduce a beta-binomial sampling model that correctly incorporates both sources of variation. We show how to fit the parameters of this model, and introduce a test statistic for differential expression similar to a two-sample t-test. Contact: [email protected] Supplementary information http://bioinformatics.mdanderson.org/ Includes Matlab and R code for fitting the model. * To whom correspondence should be addressed.
Keith A. Baggerly, Jeffrey S. Morris, C. Marcelo Aldaz
Bioinform.1