EDBT 2026 Demo / reviewers in the wild / expert
Bobbie-Jo M. Webb-Robertson
dblp:99/628
· DBLP profile ↗
15ranked-venue papers
8as first author
0since 2021 · last 2013
0000-0002-4744-2397ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 6 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 100% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
proteomics |
0.5 | 5 | 2011 | Improved quality control processing of peptide-centric LC-MS proteomics data · Bioinform. 2011 A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2010 A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2008 |
Bioinformatics and computational biology › proteomics › computational proteomics
proteotypic peptide prediction |
0.2 | 2 | 2010 | A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2010 A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2008 |
Bioinformatics and computational biology › structural biology
protein structure and function |
0.1 | 1 | 2008 | SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detection · Bioinform. 2008 |
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection |
0.1 | 1 | 2008 | SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detection · Bioinform. 2008 |
Bioinformatics and computational biology › proteomics
peptide identification |
0.1 | 1 | 2006 | Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006 |
Visualization and visual analytics
biological data visualization |
0.1 | 1 | 2006 | Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.1 | 2 | 2010 | A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2010 A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics · Bioinform. 2008 |
Bioinformatics and computational biology
multi-omics data integration |
0.0 | 1 | 2010 | VIBE 2.0: Visual Integration for Bayesian Evaluation · Bioinform. 2010 |
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search |
0.0 | 1 | 2008 | SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detection · Bioinform. 2008 |
Bioinformatics and computational biology
genomics |
0.0 | 1 | 2007 | PQuad - a visual analysis platform for proteomic data exploration of microbial organisms · Bioinform. 2007 |
Bioinformatics and computational biology › genome annotation
prokaryotic genome annotation |
0.0 | 1 | 2007 | PQuad - a visual analysis platform for proteomic data exploration of microbial organisms · Bioinform. 2007 |
High-performance computing
high-throughput computing |
0.0 | 1 | 2006 | Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.3simulation · 0.1multivariate statistics · 0.1interactive visualization · 0.1bayesian data fusion · 0.1statistical sampling · 0.1semi-supervised learning · 0.1multiresolution visualization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Assessing the quality of bioforensic signaturesabstractWe present a mathematical framework for assessing the quality of signature systems in terms of fidelity, risk, cost, other attributes, and utility-a method we call Signature Quality Metrics (SQM). We demonstrate the SQM approach by assessing the quality of a signature system designed to predict the culture medium used to grow a microorganism. The system consists of four chemical assays and a Bayesian network that estimates the probabilities the microorganism was grown using one of eleven culture media. We evaluated fifteen combinations of the signature system by removing one or more of the assays from the Bayes net. We show how SQM can be used to compare the various combinations while accounting for the tradeoffs among three attributes of interest: fidelity, cost, and the amount of sample material consumed by the assays. Landon H. Sego, Aimee E. Holmes, Luke J. Gosink, Bobbie-Jo M. Webb-Robertson, Helen W. Kreuzer, Richard M. Anderson, Alan J. Brothers, Courtney D. Corley, Mark R. Tardiff |
ISI | 4 |
| 2011 | Improved quality control processing of peptide-centric LC-MS proteomics dataabstractMOTIVATION: In the analysis of differential peptide peak intensities (i.e. abundance measures), LC-MS analyses with poor quality peptide abundance data can bias downstream statistical analyses and hence the biological interpretation for an otherwise high-quality dataset. Although considerable effort has been placed on assuring the quality of the peptide identification with respect to spectral processing, to date quality assessment of the subsequent peptide abundance data matrix has been limited to a subjective visual inspection of run-by-run correlation or individual peptide components. Identifying statistical outliers is a critical step in the processing of proteomics data as many of the downstream statistical analyses [e.g. analysis of variance (ANOVA)] rely upon accurate estimates of sample variance, and their results are influenced by extreme values. RESULTS: We describe a novel multivariate statistical strategy for the identification of LC-MS runs with extreme peptide abundance distributions. Comparison with current method (run-by-run correlation) demonstrates a significantly better rate of identification of outlier runs by the multivariate strategy. Simulation studies also suggest that this strategy significantly outperforms correlation alone in the identification of statistically extreme liquid chromatography-mass spectrometry (LC-MS) runs. AVAILABILITY: https://www.biopilot.org/docs/Software/RMD.php CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online. Melissa M. Matzke, Katrina M. Waters, Thomas O. Metz, Jon M. Jacobs, Amy C. Sims, Ralph S. Baric, Joel G. Pounds, Bobbie-Jo M. Webb-Robertson |
Bioinform. | 8 |
| 2010 | VIBE 2.0: Visual Integration for Bayesian EvaluationabstractUNLABELLED: Data fusion methods are powerful tools for evaluating experiments designed to discover measurable features of directly unobservable systems. We describe an interactive software platform, Visual Integration for Bayesian Evaluation, that ingests or creates bayesian posterior probability matrices, performs data fusion and allows the user to interactively evaluate the classification power of fusing various combinations of data sources, such as transcriptomic, proteomics, metabolomics, biochemistry and function. AVAILABILITY: http://omics.pnl.gov/software/VIBE.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nathaniel Beagley, Kelly G. Stratton, Bobbie-Jo M. Webb-Robertson |
Bioinform. | 3 |
| 2010 | A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomicsabstractAbstract Motivation: The standard approach to identifying peptides based on accurate mass and elution time (AMT) compares profiles obtained from a high resolution mass spectrometer to a database of peptides previously identified from tandem mass spectrometry (MS/MS) studies. It would be advantageous, with respect to both accuracy and cost, to only search for those peptides that are detectable by MS (proteotypic). Results: We present a support vector machine (SVM) model that uses a simple descriptor space based on 35 properties of amino acid content, charge, hydrophilicity and polarity for the quantitative prediction of proteotypic peptides. Using three independently derived AMT databases (Shewanella oneidensis, Salmonella typhimurium, Yersinia pestis) for training and validation within and across species, the SVM resulted in an average accuracy measure of ∼0.83 with an SD of <0.038. Furthermore, we demonstrate that these results are achievable with a small set of 13 variables and can achieve high proteome coverage. Availability: http://omics.pnl.gov/software/STEPP.php Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Bobbie-Jo M. Webb-Robertson, William R. Cannon, Christopher S. Oehmen, Anuj R. Shah, Vidhya Gurumoorthi, Mary S. Lipton, Katrina M. Waters |
Bioinform. | 1 |
| 2010 | Physicochemical property distributions for accurate and rapid pairwise protein homology detectionabstractBACKGROUND: The challenge of remote homology detection is that many evolutionarily related sequences have very little similarity at the amino acid level. Kernel-based discriminative methods, such as support vector machines (SVMs), that use vector representations of sequences derived from sequence properties have been shown to have superior accuracy when compared to traditional approaches for the task of remote homology detection. RESULTS: We introduce a new method for feature vector representation based on the physicochemical properties of the primary protein sequence. A distribution of physicochemical property scores are assembled from 4-mers of the sequence and normalized based on the null distribution of the property over all possible 4-mers. With this approach there is little computational cost associated with the transformation of the protein into feature space, and overall performance in terms of remote homology detection is comparable with current state-of-the-art methods. We demonstrate that the features can be used for the task of pairwise remote homology detection with improved accuracy versus sequence-based methods such as BLAST and other feature-based methods of similar computational cost. CONCLUSIONS: A protein feature method based on physicochemical properties is a viable approach for extracting features in a computationally inexpensive manner while retaining the sensitivity of SVM protein homology detection. Furthermore, identifying features that can be used for generic pairwise homology detection in lieu of family-based homology detection is important for applications such as large database searches and comparative genomics. Bobbie-Jo M. Webb-Robertson, Kyle G. Ratuiste, Christopher S. Oehmen |
BMC Bioinform. | 1 |
| 2008 | Dimension Reduction via Unsupervised Learning Yields Significant Computational Improvements for Support Vector Machine Based Protein Family ClassificationabstractReducing the dimension of vectors used in training support vector machines (SVMs) results in a proportional speedup in training time. For large-scale problems this can make the difference between tractable and intractable training tasks. However, it is critical that classifiers trained on reduced datasets perform as reliably as their counterparts trained on high-dimensional data. We assessed principal component analysis (PCA) and sequential project pursuit (SPP) as dimension reduction strategies in the biology application of classifying proteins into well-defined functional ‘families’ (SVM-based protein family classification) by their impact on run-time, sensitivity and selectivity. Homology vectors of 4352 elements were reduced to approximately 2% of the original data size using PCA and SPP without significantly affecting accuracy, while leading to approximately a 28-fold speedup in run-time. Bobbie-Jo M. Webb-Robertson, Melissa M. Matzke, Christopher S. Oehmen |
ICMLA | 1 |
| 2008 | SVM-HUSTLE - an iterative semi-supervised machine learning approach for pairwise protein remote homology detectionabstractMOTIVATION: As the amount of biological sequence data continues to grow exponentially we face the increasing challenge of assigning function to this enormous molecular 'parts list'. The most popular approaches to this challenge make use of the simplifying assumption that similar functional molecules, or proteins, sometimes have similar composition, or sequence. However, these algorithms often fail to identify remote homologs (proteins with similar function but dissimilar sequence) which often are a significant fraction of the total homolog collection for a given sequence. We introduce a Support Vector Machine (SVM)-based tool to detect homology using semi-supervised iterative learning (SVM-HUSTLE) that identifies significantly more remote homologs than current state-of-the-art sequence or cluster-based methods. As opposed to building profiles or position specific scoring matrices, SVM-HUSTLE builds an SVM classifier for a query sequence by training on a collection of representative high-confidence training sets, recruits additional sequences and assigns a statistical measure of homology between a pair of sequences. SVM-HUSTLE combines principles of semi-supervised learning theory with statistical sampling to create many concurrent classifiers to iteratively detect and refine, on-the-fly, patterns indicating homology. RESULTS: When compared against existing methods for identifying protein homologs (BLAST, PSI-BLAST, COMPASS, PROF_SIM, RANKPROP and their variants) on two different benchmark datasets SVM-HUSTLE significantly outperforms each of the above methods using the most stringent ROC(1) statistic with P-values less than 1e-20. SVM-HUSTLE also yields results comparable to HHSearch but at a substantially reduced computational cost since we do not require the construction of HMMs. AVAILABILITY: The software executable to run SVM-HUSTLE can be downloaded from http://www.sysbio.org/sysbio/networkbio/svm_hustle Anuj R. Shah, Christopher S. Oehmen, Bobbie-Jo M. Webb-Robertson |
Bioinform. | 3 |
| 2008 | A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomicsabstractMOTIVATION: The standard approach to identifying peptides based on accurate mass and elution time (AMT) compares profiles obtained from a high resolution mass spectrometer to a database of peptides previously identified from tandem mass spectrometry (MS/MS) studies. It would be advantageous, with respect to both accuracy and cost, to only search for those peptides that are detectable by MS (proteotypic). RESULTS: We present a support vector machine (SVM) model that uses a simple descriptor space based on 35 properties of amino acid content, charge, hydrophilicity and polarity for the quantitative prediction of proteotypic peptides. Using three independently derived AMT databases (Shewanella oneidensis, Salmonella typhimurium, Yersinia pestis) for training and validation within and across species, the SVM resulted in an average accuracy measure of 0.8 with a SD of <0.025. Furthermore, we demonstrate that these results are achievable with a small set of 12 variables and can achieve high proteome coverage. AVAILABILITY: http://omics.pnl.gov/software/STEPP.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bobbie-Jo M. Webb-Robertson, William R. Cannon, Christopher S. Oehmen, Anuj R. Shah, Vidhya Gurumoorthi, Mary S. Lipton, Katrina M. Waters |
Bioinform. | 1 |
| 2008 | Pairwise covariance adds little to secondary structure prediction but improves the prediction of non-canonical local structureabstractBACKGROUND: Amino acid sequence probability distributions, or profiles, have been used successfully to predict secondary structure and local structure in proteins. Profile models assume the statistical independence of each position in the sequence, but the energetics of protein folding is better captured in a scoring function that is based on pairwise interactions, like a force field. RESULTS: I-sites motifs are short sequence/structure motifs that populate the protein structure database due to energy-driven convergent evolution. Here we show that a pairwise covariant sequence model does not predict alpha helix or beta strand significantly better overall than a profile-based model, but it does improve the prediction of certain loop motifs. The finding is best explained by considering secondary structure profiles as multivariant, all-or-none models, which subsume covariant models. Pairwise covariance is nonetheless present and energetically rational. Examples of negative design are present, where the covariances disfavor non-native structures. CONCLUSION: Measured pairwise covariances are shown to be statistically robust in cross-validation tests, as long as the amino acid alphabet is reduced to nine classes. An updated I-sites local structure motif library that provides sequence covariance information for all types of local structure in globular proteins and a web server for local structure prediction are available at http://www.bioinfo.rpi.edu/bystrc/hmmstr/server.php. Christopher Bystroff, Bobbie-Jo M. Webb-Robertson |
BMC Bioinform. | 2 |
| 2008 | Measuring Global Credibility with Application to Local Sequence AlignmentabstractComputational biology is replete with high-dimensional (high-D) discrete prediction and inference problems, including sequence alignment, RNA structure prediction, phylogenetic inference, motif finding, prediction of pathways, and model selection problems in statistical genetics. Even though prediction and inference in these settings are uncertain, little attention has been focused on the development of global measures of uncertainty. Regardless of the procedure employed to produce a prediction, when a procedure delivers a single answer, that answer is a point estimate selected from the solution ensemble, the set of all possible solutions. For high-D discrete space, these ensembles are immense, and thus there is considerable uncertainty. We recommend the use of Bayesian credibility limits to describe this uncertainty, where a (1-alpha)%, 0< or =alpha< or =1, credibility limit is the minimum Hamming distance radius of a hyper-sphere containing (1-alpha)% of the posterior distribution. Because sequence alignment is arguably the most extensively used procedure in computational biology, we employ it here to make these general concepts more concrete. The maximum similarity estimator (i.e., the alignment that maximizes the likelihood) and the centroid estimator (i.e., the alignment that minimizes the mean Hamming distance from the posterior weighted ensemble of alignments) are used to demonstrate the application of Bayesian credibility limits to alignment estimators. Application of Bayesian credibility limits to the alignment of 20 human/rodent orthologous sequence pairs and 125 orthologous sequence pairs from six Shewanella species shows that credibility limits of the alignments of promoter sequences of these species vary widely, and that centroid alignments dependably have tighter credibility limits than traditional maximum similarity alignments. Bobbie-Jo M. Webb-Robertson, Lee Ann McCue, Charles E. Lawrence |
PLoS Comput. Biol. | 1 |
| 2007 | Support Vector Machine Classification of Probability Models and Peptide Features for Improved Peptide Identification from Shotgun ProteomicsabstractMass spectrometry (MS)-based proteomics is a powerful and popular high-throughput process for characterizing the global protein content of a sample. In shotgun proteomics, typically proteins are digested into fragments (peptides) prior to mass analysis, and the presence of a protein in inferred from the identification of its constituent peptides. Thus, accurate proteome characterization is dependent upon the accuracy of this peptide identification step. Database search routines generate predicted spectra for all peptides derived from the known genome information, and thus, identify a peptide by 'matching' an experimental to a predicted spectrum. However, due to many problems, such as incomplete fragmentation, this process results in a large number of false positives. We present a new scoring algorithm that integrates probabilistic database scoring metrics (from the MSPolygraph program) with physico-chemical properties in a support vector machine (SVM). We demonstrate that this peptide identification classifier SVM (PICS) score is not only more accurate than the single best database scoring metric, but is also significantly more accurate than models derived using a linear discriminant analysis, decision tree, or artificial neural network. Bobbie-Jo M. Webb-Robertson, Christopher S. Oehmen, William R. Cannon |
ICMLA | 1 |
| 2007 | Current trends in computational inference from mass spectrometry-based proteomicsabstractMass spectrometry offers a high-throughput approach to quantifying the proteome associated with a biological sample and hence has become the primary approach of proteomic analyses. Computation is tightly coupled to this advanced technological platform as a required component of not only peptide and protein identification, but quantification and functional inference, such as protein modifications and interactions. Proteomics faces several key computational challenges such as identification of proteins and peptides from tandem mass spectra as well as their quantitation. In addition, the application of proteomics to systems biology requires understanding the functional proteome, including how the dynamics of the cell change in response to protein modifications and complex interactions between biomolecules. This review presents an overview of recently developed methods and their impact on these core computational challenges currently facing proteomics. Bobbie-Jo M. Webb-Robertson, William R. Cannon |
Briefings Bioinform. | 1 |
| 2007 | PQuad - a visual analysis platform for proteomic data exploration of microbial organismsabstractUNLABELLED: The visual Platform for Proteomics Peptide and Protein data exploration (PQuad) is a multi-resolution environment that visually integrates genomic and proteomic data for prokaryotic systems, overlays categorical annotation and compares differential expression experiments. PQuad requires Java 1.5 and has been tested to run across different operating systems. AVAILABILITY: http://ncrr.pnl.gov/software. Bobbie-Jo M. Webb-Robertson, Elena S. Peterson, Mudita Singhal, Kyle R. Klicker, Christopher S. Oehmen, Joshua N. Adkins, Susan L. Havre |
Bioinform. | 1 |
| 2006 | Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledgeabstractFor the SC|06 analytics challenge, we demonstrate an end-to-end solution for processing data produced by high-throughput mass spectrometry (MS)-based proteomics so biological hypotheses can be explored. This approach is based on a tool called the Bioinformatics Resource Manager (BRM) which will interact with high-performance architecture and experimental data sources to provide high-throughput analytics to a specific experimental dataset. Peptide identification was achieved by a high-performance code, Polygraph, which has been shown to scale well beyond 1000 processors. Visual analytics applications such as PQuad, Cytoscape, or others may be used to visualize protein identities in the context of pathways using data from public repositories such as Kyoto Encyclopedia of Genes and Genomes (KEGG). The end result was that a user can go from experimental spectra to pathway data in a single workflow reducing time-to-solution for analyzing biological data from weeks to minutes. Christopher S. Oehmen, Lee Ann McCue, Joshua N. Adkins, Katrina M. Waters, Tim Carlson, William R. Cannon, Bobbie-Jo M. Webb-Robertson, Douglas J. Baxter, Elena S. Peterson, Mudita Singhal, Anuj R. Shah, Kyle R. Klicker |
SC | 7 |
| 2004 | PQuad: Visualization of Predicted Peptides and Proteins
Susan L. Havre, Mudita Singhal, Deborah A. Payne, Bobbie-Jo M. Webb-Robertson |
IEEE Visualization | 4 |