Bobbie-Jo M. Webb-Robertson

dblp:99/628 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
0since 2021 · last 2013
0000-0002-4744-2397ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 6 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
7 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
proteomics
0.552011
Improved quality control processing of peptide-centric LC-MS proteomics data · Bioinform. 2011
A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2010
A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2008
Bioinformatics and computational biology › proteomics › computational proteomics
proteotypic peptide prediction
0.222010
A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2010
A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2008
Bioinformatics and computational biology › structural biology
protein structure and function
0.112008
SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detection · Bioinform. 2008
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection
0.112008
SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detection · Bioinform. 2008
Bioinformatics and computational biology › proteomics
peptide identification
0.112006
Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006
Visualization and visual analytics
biological data visualization
0.112006
Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.122010
A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2010
A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2008
Bioinformatics and computational biology
multi-omics data integration
0.012010
VIBE 2.0: Visual Integration for Bayesian Evaluation · Bioinform. 2010
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.012008
SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detection · Bioinform. 2008
Bioinformatics and computational biology
genomics
0.012007
PQuad - a visual analysis platform for proteomic data exploration of microbial organisms · Bioinform. 2007
Bioinformatics and computational biology › genome annotation
prokaryotic genome annotation
0.012007
PQuad - a visual analysis platform for proteomic data exploration of microbial organisms · Bioinform. 2007
High-performance computing
high-throughput computing
0.012006
Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.3simulation · 0.1multivariate statistics · 0.1interactive visualization · 0.1bayesian data fusion · 0.1statistical sampling · 0.1semi-supervised learning · 0.1multiresolution visualization · 0.1
YearPublicationVenuePosition
2013 Assessing the quality of bioforensic signatures
abstract
We present a mathematical framework for assessing the quality of signature systems in terms of fidelity, risk, cost, other attributes, and utility-a method we call Signature Quality Metrics (SQM). We demonstrate the SQM approach by assessing the quality of a signature system designed to predict the culture medium used to grow a microorganism. The system consists of four chemical assays and a Bayesian network that estimates the probabilities the microorganism was grown using one of eleven culture media. We evaluated fifteen combinations of the signature system by removing one or more of the assays from the Bayes net. We show how SQM can be used to compare the various combinations while accounting for the tradeoffs among three attributes of interest: fidelity, cost, and the amount of sample material consumed by the assays.
Landon H. Sego, Aimee E. Holmes, Luke J. Gosink, Bobbie-Jo M. Webb-Robertson, Helen W. Kreuzer, Richard M. Anderson, Alan J. Brothers, Courtney D. Corley, Mark R. Tardiff
ISI4
2011 Improved quality control processing of peptide-centric LC-MS proteomics data
abstract
MOTIVATION: In the analysis of differential peptide peak intensities (i.e. abundance measures), LC-MS analyses with poor quality peptide abundance data can bias downstream statistical analyses and hence the biological interpretation for an otherwise high-quality dataset. Although considerable effort has been placed on assuring the quality of the peptide identification with respect to spectral processing, to date quality assessment of the subsequent peptide abundance data matrix has been limited to a subjective visual inspection of run-by-run correlation or individual peptide components. Identifying statistical outliers is a critical step in the processing of proteomics data as many of the downstream statistical analyses [e.g. analysis of variance (ANOVA)] rely upon accurate estimates of sample variance, and their results are influenced by extreme values. RESULTS: We describe a novel multivariate statistical strategy for the identification of LC-MS runs with extreme peptide abundance distributions. Comparison with current method (run-by-run correlation) demonstrates a significantly better rate of identification of outlier runs by the multivariate strategy. Simulation studies also suggest that this strategy significantly outperforms correlation alone in the identification of statistically extreme liquid chromatography-mass spectrometry (LC-MS) runs. AVAILABILITY: https://www.biopilot.org/docs/Software/RMD.php CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.
Melissa M. Matzke, Katrina M. Waters, Thomas O. Metz, Jon M. Jacobs, Amy C. Sims, Ralph S. Baric, Joel G. Pounds, Bobbie-Jo M. Webb-Robertson
Bioinform.8
2010 VIBE 2.0: Visual Integration for Bayesian Evaluation
abstract
UNLABELLED: Data fusion methods are powerful tools for evaluating experiments designed to discover measurable features of directly unobservable systems. We describe an interactive software platform, Visual Integration for Bayesian Evaluation, that ingests or creates bayesian posterior probability matrices, performs data fusion and allows the user to interactively evaluate the classification power of fusing various combinations of data sources, such as transcriptomic, proteomics, metabolomics, biochemistry and function. AVAILABILITY: http://omics.pnl.gov/software/VIBE.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nathaniel Beagley, Kelly G. Stratton, Bobbie-Jo M. Webb-Robertson
Bioinform.3
2010 A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics
abstract
Abstract Motivation: The standard approach to identifying peptides based on accurate mass and elution time (AMT) compares profiles obtained from a high resolution mass spectrometer to a database of peptides previously identified from tandem mass spectrometry (MS/MS) studies. It would be advantageous, with respect to both accuracy and cost, to only search for those peptides that are detectable by MS (proteotypic). Results: We present a support vector machine (SVM) model that uses a simple descriptor space based on 35 properties of amino acid content, charge, hydrophilicity and polarity for the quantitative prediction of proteotypic peptides. Using three independently derived AMT databases (Shewanella oneidensis, Salmonella typhimurium, Yersinia pestis) for training and validation within and across species, the SVM resulted in an average accuracy measure of ∼0.83 with an SD of <0.038. Furthermore, we demonstrate that these results are achievable with a small set of 13 variables and can achieve high proteome coverage. Availability: http://omics.pnl.gov/software/STEPP.php Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Bobbie-Jo M. Webb-Robertson, William R. Cannon, Christopher S. Oehmen, Anuj R. Shah, Vidhya Gurumoorthi, Mary S. Lipton, Katrina M. Waters
Bioinform.1
2010 Physicochemical property distributions for accurate and rapid pairwise protein homology detection
abstract
BACKGROUND: The challenge of remote homology detection is that many evolutionarily related sequences have very little similarity at the amino acid level. Kernel-based discriminative methods, such as support vector machines (SVMs), that use vector representations of sequences derived from sequence properties have been shown to have superior accuracy when compared to traditional approaches for the task of remote homology detection. RESULTS: We introduce a new method for feature vector representation based on the physicochemical properties of the primary protein sequence. A distribution of physicochemical property scores are assembled from 4-mers of the sequence and normalized based on the null distribution of the property over all possible 4-mers. With this approach there is little computational cost associated with the transformation of the protein into feature space, and overall performance in terms of remote homology detection is comparable with current state-of-the-art methods. We demonstrate that the features can be used for the task of pairwise remote homology detection with improved accuracy versus sequence-based methods such as BLAST and other feature-based methods of similar computational cost. CONCLUSIONS: A protein feature method based on physicochemical properties is a viable approach for extracting features in a computationally inexpensive manner while retaining the sensitivity of SVM protein homology detection. Furthermore, identifying features that can be used for generic pairwise homology detection in lieu of family-based homology detection is important for applications such as large database searches and comparative genomics.
Bobbie-Jo M. Webb-Robertson, Kyle G. Ratuiste, Christopher S. Oehmen
BMC Bioinform.1
2008 Dimension Reduction via Unsupervised Learning Yields Significant Computational Improvements for Support Vector Machine Based Protein Family Classification
abstract
Reducing the dimension of vectors used in training support vector machines (SVMs) results in a proportional speedup in training time. For large-scale problems this can make the difference between tractable and intractable training tasks. However, it is critical that classifiers trained on reduced datasets perform as reliably as their counterparts trained on high-dimensional data. We assessed principal component analysis (PCA) and sequential project pursuit (SPP) as dimension reduction strategies in the biology application of classifying proteins into well-defined functional ‘families’ (SVM-based protein family classification) by their impact on run-time, sensitivity and selectivity. Homology vectors of 4352 elements were reduced to approximately 2% of the original data size using PCA and SPP without significantly affecting accuracy, while leading to approximately a 28-fold speedup in run-time.
Bobbie-Jo M. Webb-Robertson, Melissa M. Matzke, Christopher S. Oehmen
ICMLA1
2008 SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detection
abstract
MOTIVATION: As the amount of biological sequence data continues to grow exponentially we face the increasing challenge of assigning function to this enormous molecular 'parts list'. The most popular approaches to this challenge make use of the simplifying assumption that similar functional molecules, or proteins, sometimes have similar composition, or sequence. However, these algorithms often fail to identify remote homologs (proteins with similar function but dissimilar sequence) which often are a significant fraction of the total homolog collection for a given sequence. We introduce a Support Vector Machine (SVM)-based tool to detect homology using semi-supervised iterative learning (SVM-HUSTLE) that identifies significantly more remote homologs than current state-of-the-art sequence or cluster-based methods. As opposed to building profiles or position specific scoring matrices, SVM-HUSTLE builds an SVM classifier for a query sequence by training on a collection of representative high-confidence training sets, recruits additional sequences and assigns a statistical measure of homology between a pair of sequences. SVM-HUSTLE combines principles of semi-supervised learning theory with statistical sampling to create many concurrent classifiers to iteratively detect and refine, on-the-fly, patterns indicating homology. RESULTS: When compared against existing methods for identifying protein homologs (BLAST, PSI-BLAST, COMPASS, PROF_SIM, RANKPROP and their variants) on two different benchmark datasets SVM-HUSTLE significantly outperforms each of the above methods using the most stringent ROC(1) statistic with P-values less than 1e-20. SVM-HUSTLE also yields results comparable to HHSearch but at a substantially reduced computational cost since we do not require the construction of HMMs. AVAILABILITY: The software executable to run SVM-HUSTLE can be downloaded from http://www.sysbio.org/sysbio/networkbio/svm_hustle
Anuj R. Shah, Christopher S. Oehmen, Bobbie-Jo M. Webb-Robertson
Bioinform.3
2008 A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics
abstract
MOTIVATION: The standard approach to identifying peptides based on accurate mass and elution time (AMT) compares profiles obtained from a high resolution mass spectrometer to a database of peptides previously identified from tandem mass spectrometry (MS/MS) studies. It would be advantageous, with respect to both accuracy and cost, to only search for those peptides that are detectable by MS (proteotypic). RESULTS: We present a support vector machine (SVM) model that uses a simple descriptor space based on 35 properties of amino acid content, charge, hydrophilicity and polarity for the quantitative prediction of proteotypic peptides. Using three independently derived AMT databases (Shewanella oneidensis, Salmonella typhimurium, Yersinia pestis) for training and validation within and across species, the SVM resulted in an average accuracy measure of 0.8 with a SD of <0.025. Furthermore, we demonstrate that these results are achievable with a small set of 12 variables and can achieve high proteome coverage. AVAILABILITY: http://omics.pnl.gov/software/STEPP.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bobbie-Jo M. Webb-Robertson, William R. Cannon, Christopher S. Oehmen, Anuj R. Shah, Vidhya Gurumoorthi, Mary S. Lipton, Katrina M. Waters
Bioinform.1
2008 Pairwise covariance adds little to secondary structure prediction but improves the prediction of non-canonical local structure
abstract
BACKGROUND: Amino acid sequence probability distributions, or profiles, have been used successfully to predict secondary structure and local structure in proteins. Profile models assume the statistical independence of each position in the sequence, but the energetics of protein folding is better captured in a scoring function that is based on pairwise interactions, like a force field. RESULTS: I-sites motifs are short sequence/structure motifs that populate the protein structure database due to energy-driven convergent evolution. Here we show that a pairwise covariant sequence model does not predict alpha helix or beta strand significantly better overall than a profile-based model, but it does improve the prediction of certain loop motifs. The finding is best explained by considering secondary structure profiles as multivariant, all-or-none models, which subsume covariant models. Pairwise covariance is nonetheless present and energetically rational. Examples of negative design are present, where the covariances disfavor non-native structures. CONCLUSION: Measured pairwise covariances are shown to be statistically robust in cross-validation tests, as long as the amino acid alphabet is reduced to nine classes. An updated I-sites local structure motif library that provides sequence covariance information for all types of local structure in globular proteins and a web server for local structure prediction are available at http://www.bioinfo.rpi.edu/bystrc/hmmstr/server.php.
Christopher Bystroff, Bobbie-Jo M. Webb-Robertson
BMC Bioinform.2
2008 Measuring Global Credibility with Application to Local Sequence Alignment
abstract
Computational biology is replete with high-dimensional (high-D) discrete prediction and inference problems, including sequence alignment, RNA structure prediction, phylogenetic inference, motif finding, prediction of pathways, and model selection problems in statistical genetics. Even though prediction and inference in these settings are uncertain, little attention has been focused on the development of global measures of uncertainty. Regardless of the procedure employed to produce a prediction, when a procedure delivers a single answer, that answer is a point estimate selected from the solution ensemble, the set of all possible solutions. For high-D discrete space, these ensembles are immense, and thus there is considerable uncertainty. We recommend the use of Bayesian credibility limits to describe this uncertainty, where a (1-alpha)%, 0< or =alpha< or =1, credibility limit is the minimum Hamming distance radius of a hyper-sphere containing (1-alpha)% of the posterior distribution. Because sequence alignment is arguably the most extensively used procedure in computational biology, we employ it here to make these general concepts more concrete. The maximum similarity estimator (i.e., the alignment that maximizes the likelihood) and the centroid estimator (i.e., the alignment that minimizes the mean Hamming distance from the posterior weighted ensemble of alignments) are used to demonstrate the application of Bayesian credibility limits to alignment estimators. Application of Bayesian credibility limits to the alignment of 20 human/rodent orthologous sequence pairs and 125 orthologous sequence pairs from six Shewanella species shows that credibility limits of the alignments of promoter sequences of these species vary widely, and that centroid alignments dependably have tighter credibility limits than traditional maximum similarity alignments.
Bobbie-Jo M. Webb-Robertson, Lee Ann McCue, Charles E. Lawrence
PLoS Comput. Biol.1
2007 Support Vector Machine Classification of Probability Models and Peptide Features for Improved Peptide Identification from Shotgun Proteomics
abstract
Mass spectrometry (MS)-based proteomics is a powerful and popular high-throughput process for characterizing the global protein content of a sample. In shotgun proteomics, typically proteins are digested into fragments (peptides) prior to mass analysis, and the presence of a protein in inferred from the identification of its constituent peptides. Thus, accurate proteome characterization is dependent upon the accuracy of this peptide identification step. Database search routines generate predicted spectra for all peptides derived from the known genome information, and thus, identify a peptide by 'matching' an experimental to a predicted spectrum. However, due to many problems, such as incomplete fragmentation, this process results in a large number of false positives. We present a new scoring algorithm that integrates probabilistic database scoring metrics (from the MSPolygraph program) with physico-chemical properties in a support vector machine (SVM). We demonstrate that this peptide identification classifier SVM (PICS) score is not only more accurate than the single best database scoring metric, but is also significantly more accurate than models derived using a linear discriminant analysis, decision tree, or artificial neural network.
Bobbie-Jo M. Webb-Robertson, Christopher S. Oehmen, William R. Cannon
ICMLA1
2007 Current trends in computational inference from mass spectrometry-based proteomics
abstract
Mass spectrometry offers a high-throughput approach to quantifying the proteome associated with a biological sample and hence has become the primary approach of proteomic analyses. Computation is tightly coupled to this advanced technological platform as a required component of not only peptide and protein identification, but quantification and functional inference, such as protein modifications and interactions. Proteomics faces several key computational challenges such as identification of proteins and peptides from tandem mass spectra as well as their quantitation. In addition, the application of proteomics to systems biology requires understanding the functional proteome, including how the dynamics of the cell change in response to protein modifications and complex interactions between biomolecules. This review presents an overview of recently developed methods and their impact on these core computational challenges currently facing proteomics.
Bobbie-Jo M. Webb-Robertson, William R. Cannon
Briefings Bioinform.1
2007 PQuad - a visual analysis platform for proteomic data exploration of microbial organisms
abstract
UNLABELLED: The visual Platform for Proteomics Peptide and Protein data exploration (PQuad) is a multi-resolution environment that visually integrates genomic and proteomic data for prokaryotic systems, overlays categorical annotation and compares differential expression experiments. PQuad requires Java 1.5 and has been tested to run across different operating systems. AVAILABILITY: http://ncrr.pnl.gov/software.
Bobbie-Jo M. Webb-Robertson, Elena S. Peterson, Mudita Singhal, Kyle R. Klicker, Christopher S. Oehmen, Joshua N. Adkins, Susan L. Havre
Bioinform.1
2006 Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge
abstract
For the SC|06 analytics challenge, we demonstrate an end-to-end solution for processing data produced by high-throughput mass spectrometry (MS)-based proteomics so biological hypotheses can be explored. This approach is based on a tool called the Bioinformatics Resource Manager (BRM) which will interact with high-performance architecture and experimental data sources to provide high-throughput analytics to a specific experimental dataset. Peptide identification was achieved by a high-performance code, Polygraph, which has been shown to scale well beyond 1000 processors. Visual analytics applications such as PQuad, Cytoscape, or others may be used to visualize protein identities in the context of pathways using data from public repositories such as Kyoto Encyclopedia of Genes and Genomes (KEGG). The end result was that a user can go from experimental spectra to pathway data in a single workflow reducing time-to-solution for analyzing biological data from weeks to minutes.
Christopher S. Oehmen, Lee Ann McCue, Joshua N. Adkins, Katrina M. Waters, Tim Carlson, William R. Cannon, Bobbie-Jo M. Webb-Robertson, Douglas J. Baxter, Elena S. Peterson, Mudita Singhal, Anuj R. Shah, Kyle R. Klicker
SC7
2004 PQuad: Visualization of Predicted Peptides and Proteins
Susan L. Havre, Mudita Singhal, Deborah A. Payne, Bobbie-Jo M. Webb-Robertson
IEEE Visualization4