VLDB 2026 Research / reviewers in the wild / expert
Robert E. Thurman
dblp:88/4639
· DBLP profile ↗
5ranked-venue papers
0as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Kernel, tree and ensemble methods · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics
genomic data compression |
0.1 | 1 | 2012 | BEDOPS: high-performance genomic feature operations · Bioinform. 2012 |
Bioinformatics and computational biology › genomics › genomic interval analysis
genomic interval operations |
0.1 | 1 | 2012 | BEDOPS: high-performance genomic feature operations · Bioinform. 2012 |
Bioinformatics and computational biology
genomics |
0.1 | 1 | 2012 | BEDOPS: high-performance genomic feature operations · Bioinform. 2012 |
Bioinformatics and computational biology › genomics › genomic data compression
lossless compression |
0.1 | 1 | 2012 | BEDOPS: high-performance genomic feature operations · Bioinform. 2012 |
Bioinformatics and computational biology › epigenomics
chromatin accessibility |
0.1 | 1 | 2008 | Automated mapping of large-scale chromatin structure in ENCODE · Bioinform. 2008 |
Bioinformatics and computational biology › epigenomics
chromatin conformation |
0.1 | 1 | 2008 | Automated mapping of large-scale chromatin structure in ENCODE · Bioinform. 2008 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › structured kernel
string kernel |
0.1 | 1 | 2005 | Kernels for gene regulatory regions · NIPS 2005 |
Methods — techniques the papers use, named apart from their topics
hidden markov model · 0.2parallel processing · 0.1interval tree · 0.1support vector machine · 0.1multiple alignment kernel · 0.1bayesian hierarchical change-point model · 0.1wavelet smoothing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | BEDOPS: high-performance genomic feature operationsabstractAbstract Summary: The large and growing number of genome-wide datasets highlights the need for high-performance feature analysis and data comparison methods, in addition to efficient data storage and retrieval techniques. We introduce BEDOPS, a software suite for common genomic analysis tasks which offers improved flexibility, scalability and execution time characteristics over previously published packages. The suite includes a utility to compress large inputs into a lossless format that can provide greater space savings and faster data extractions than alternatives. Availability: http://code.google.com/p/bedops/ includes binaries, source and documentation. Contact: [email protected] and [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Shane J. Neph, Scott Kuehn, Alex P. Reynolds, Eric Haugen, Robert E. Thurman, Audra K. Johnson, Eric Rynes, Matthew T. Maurano, Jeff Vierstra, Sean Thomas, Richard S. Sandstrom, Richard Humbert, John A. Stamatoyannopoulos |
Bioinform. | 5 |
| 2008 | Automated mapping of large-scale chromatin structure in ENCODEabstractMOTIVATION: A recently developed DNaseI assay has given us our first genome-wide view of chromatin structure. In addition to cataloging DNaseI hypersensitive sites, these data allows us to more completely characterize overall features of chromatin accessibility. We employed a Bayesian hierarchical change-point model (CPM), a generalization of a hidden Markov Model (HMM), to characterize tiled microarray DNaseI sensitivity data available from the ENCODE project. RESULTS: Our analysis shows that the accessibility of chromatin to cleavage by DNaseI is well described by a four state model of local segments with each state described by a continuous mixture of Gaussian variables. The CPM produces a better fit to the observed data than the HMM. The large posterior probability for the four-state CPM suggests that the data falls naturally into four classes of regions, which we call major and minor DNaseI hypersensitive sites (DHSs), regions of intermediate sensitivity, and insensitive regions. These classes agree well with a model of chromatin in which local disruptions (DHSs) are concentrated within larger domains of intermediate sensitivity, the accessibility islands. The CPM assigns 92% of the bases within the ENCODE regions to the insensitive regions. The 5.8% of the bases that are in regions of intermediate sensitivity are clearly enriched in functional elements, including genes and activating histone modifications, while the remaining 2.2% of the bases in hypersensitive regions are very strongly enriched in these elements. AVAILABILITY: The CPM software is available upon request from the authors. Heng Lian 0002, William A. Thompson, Robert E. Thurman, John A. Stamatoyannopoulos, William Stafford Noble, Charles E. Lawrence |
Bioinform. | 3 |
| 2008 | Predicting Human Nucleosome Occupancy from Primary SequenceabstractNucleosomes are the fundamental repeating unit of chromatin and comprise the structural building blocks of the living eukaryotic genome. Micrococcal nuclease (MNase) has long been used to delineate nucleosomal organization. Microarray-based nucleosome mapping experiments in yeast chromatin have revealed regularly-spaced translational phasing of nucleosomes. These data have been used to train computational models of sequence-directed nuclesosome positioning, which have identified ubiquitous strong intrinsic nucleosome positioning signals. Here, we successfully apply this approach to nucleosome positioning experiments from human chromatin. The predictions made by the human-trained and yeast-trained models are strongly correlated, suggesting a shared mechanism for sequence-based determination of nucleosome occupancy. In addition, we observed striking complementarity between classifiers trained on experimental data from weakly versus heavily digested MNase samples. In the former case, the resulting model accurately identifies nucleosome-forming sequences; in the latter, the classifier excels at identifying nucleosome-free regions. Using this model we are able to identify several characteristics of nucleosome-forming and nucleosome-disfavoring sequences. First, by combining results from each classifier applied de novo across the human ENCODE regions, the classifier reveals distinct sequence composition and periodicity features of nucleosome-forming and nucleosome-disfavoring sequences. Short runs of dinucleotide repeat appear as a hallmark of nucleosome-disfavoring sequences, while nucleosome-forming sequences contain short periodic runs of GC base pairs. Second, we show that nucleosome phasing is most frequently predicted flanking nucleosome-free regions. The results suggest that the major mechanism of nucleosome positioning in vivo is boundary-event-driven and affirm the classical statistical positioning theory of nucleosome organization. Shobhit Gupta, Jonathan H. Dennis, Robert E. Thurman, Robert E. Kingston, John A. Stamatoyannopoulos, William Stafford Noble |
PLoS Comput. Biol. | 3 |
| 2007 | Unsupervised segmentation of continuous genomic dataabstractUNLABELLED: The advent of high-density, high-volume genomic data has created the need for tools to summarize large datasets at multiple scales. HMMSeg is a command-line utility for the scale-specific segmentation of continuous genomic data using hidden Markov models (HMMs). Scale specificity is achieved by an optional wavelet-based smoothing operation. HMMSeg is capable of handling multiple datasets simultaneously, rendering it ideal for integrative analysis of expression, phylogenetic and functional genomic data. AVAILABILITY: http://noble.gs.washington.edu/proj/hmmseg Nathan Day, Andrew Hemmaplardh, Robert E. Thurman, John A. Stamatoyannopoulos, William Stafford Noble |
Bioinform. | 3 |
| 2005 | Kernels for gene regulatory regionsabstractWe describe a hierarchy of motif-based kernels for multiple alignments of biological sequences, particularly suitable to process regulatory regions of genes. The kernels incorporate progressively more information, with the most complex kernel accounting for a multiple alignment of orthologous regions, the phylogenetic tree relating the species, and the prior knowledge that relevant sequence patterns occur in conserved motif blocks. These kernels can be used in the presence of a library of known transcription factor binding sites, or de novo by iterating over all k -mers of a given length. In the latter mode, a discriminative classifier built from such a kernel not only recognizes a given class of promoter regions, but as a side effect simultaneously identifies a collection of relevant, discriminative sequence motifs. We demonstrate the utility of the motif-based multiple alignment kernels by using a collection of aligned promoter regions from five yeast species to recognize classes of cell-cycle regulated genes. Supplementary data is available at http://noble.gs.washington.edu/proj/pkernel. Jean-Philippe Vert, Robert E. Thurman, William Stafford Noble |
NIPS | 2 |