Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Adam C. Siepel

dblp:s/AdamCSiepel · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
1since 2021 · last 2021
0000-0002-3557-7219ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
RNA sequencing
0.512021
Deconvolution of expression for nascent RNA-sequencing data (DENR) highlights pre-RNA isoform diversity in human cells · Bioinform. 2021
Bioinformatics and computational biology
comparative genomics
0.422019
PhastWeb: a web interface for evolutionary conservation scoring of multiple sequence alignments using phastCons and phyloP · Bioinform. 2019
An algorithm to enumerate all sorting reversals · RECOMB 2002
Bioinformatics and computational biology
multiple sequence alignment
0.412019
PhastWeb: a web interface for evolutionary conservation scoring of multiple sequence alignments using phastCons and phyloP · Bioinform. 2019
Bioinformatics and computational biology
molecular evolution
0.232006
New Methods for Detecting Lineage-Specific Selection · RECOMB 2006
Computational identification of evolutionarily conserved exons · RECOMB 2004
Combining phylogenetic and hidden Markov models in biosequence analysis · RECOMB 2003
Bioinformatics and computational biology › phylogenetics › statistical phylogenetics
phylogenetic hidden markov model
0.122004
Computational identification of evolutionarily conserved exons · RECOMB 2004
Combining phylogenetic and hidden Markov models in biosequence analysis · RECOMB 2003
Bioinformatics and computational biology › phylogenetics
evolutionary history reconstruction
0.112008
Reconstructing the Evolutionary History of Complex Human Gene Clusters · RECOMB 2008
Bioinformatics and computational biology
phylogenetics
0.112008
Reconstructing the Evolutionary History of Complex Human Gene Clusters · RECOMB 2008
Bioinformatics and computational biology › genome annotation
gene prediction
0.012004
Computational identification of evolutionarily conserved exons · RECOMB 2004
Bioinformatics and computational biology › sequence analysis
biosequence analysis
0.012003
Combining phylogenetic and hidden Markov models in biosequence analysis · RECOMB 2003
Bioinformatics and computational biology › molecular evolution
phylogenetic model
0.012003
Combining phylogenetic and hidden Markov models in biosequence analysis · RECOMB 2003
Bioinformatics and computational biology
sequence analysis
0.012003
Combining phylogenetic and hidden Markov models in biosequence analysis · RECOMB 2003
Bioinformatics and computational biology › comparative genomics
genome rearrangement
0.012002
An algorithm to enumerate all sorting reversals · RECOMB 2002
Bioinformatics and computational biology › data integration
bioinformatics resource integration
0.012001
ISYS: a decentralized, component-based approach to the integration of heterogeneous bioinformatics resources · Bioinform. 2001

Methods — techniques the papers use, named apart from their topics

transcription start site prediction · 0.5mixture model · 0.5machine learning · 0.5phylop · 0.4phylogenetic analysis · 0.4phastcons · 0.4phylogenetic reconstruction · 0.1statistical test · 0.1indel modeling · 0.0context-dependent phylogenetic model · 0.0
YearPublicationVenuePosition
2021 Deconvolution of expression for nascent RNA-sequencing data (DENR) highlights pre-RNA isoform diversity in human cells
abstract
MOTIVATION: Quantification of isoform abundance has been extensively studied at the mature RNA level using RNA-seq but not at the level of precursor RNAs using nascent RNA sequencing. RESULTS: We address this problem with a new computational method called Deconvolution of Expression for Nascent RNA-sequencing data (DENR), which models nascent RNA-sequencing read-counts as a mixture of user-provided isoforms. The baseline algorithm is enhanced by machine-learning predictions of active transcription start sites and an adjustment for the typical 'shape profile' of read-counts along a transcription unit. We show that DENR outperforms simple read-count-based methods for estimating gene and isoform abundances, and that transcription of multiple pre-RNA isoforms per gene is widespread, with frequent differences between cell types. In addition, we provide evidence that a majority of human isoform diversity derives from primary transcription rather than from post-transcriptional processes. AVAILABILITY AND IMPLEMENTATION: DENR and nascentRNASim are freely available at https://github.com/CshlSiepelLab/DENR (version v1.0.0) and https://github.com/CshlSiepelLab/nascentRNASim (version v0.3.0). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yixin Zhao, Noah Dukler, Gilad Barshad, Shushan Toneyan, Charles G. Danko, Adam C. Siepel
Bioinform.6
2019 PhastWeb: a web interface for evolutionary conservation scoring of multiple sequence alignments using phastCons and phyloP
abstract
SUMMARY: The Phylogenetic Analysis with Space/Time models (PHAST) package is a widely used software package for comparative genomics that has been freely available for download since 2002. Here, we introduce a web interface (phastWeb) that makes it possible to use two of the most popular programs in PHAST, phastCons and phyloP, without downloading and installing the PHAST software. This interface allows users to upload a sequence alignment and either upload a corresponding phylogeny or have one estimated from the alignment. After processing, users can visualize alignments and conservation scores as genome browser tracks and download estimated tree models and raw scores for further analysis. Altogether, this resource makes key features of the PHAST package conveniently available to a broad audience. AVAILABILITY AND IMPLEMENTATION: PhastWeb is freely available on the web at http://compgen.cshl.edu/phastweb/. The website provides instructions as well as examples.
Ritika Ramani, Katie Krumholz, Yifei Huang 0004, Adam C. Siepel
Bioinform.4
2014 Enabling large-scale next-generation sequence assembly with Blacklight
abstract
A variety of extremely challenging biological sequence analyses were conducted on the XSEDE large shared memory resource Blacklight, using current bioinformatics tools and encompassing a wide range of scientific applications. These include genomic sequence assembly, very large metagenomic sequence assembly, transcriptome assembly, and sequencing error correction. The data sets used in these analyses included uncategorized fungal species, reference microbial data, very large soil and human gut microbiome sequence data, and primate transcriptomes, composed of both short-read and long-read sequence data. A new parallel command execution program was developed on the Blacklight resource to handle some of these analyses. These results, initially reported previously at XSEDE13 and expanded here, represent significant advances for their respective scientific communities. The breadth and depth of the results achieved demonstrate the ease of use, versatility, and unique capabilities of the Blacklight XSEDE resource for scientific analysis of genomic and transcriptomic sequence data, and the power of these resources, together with XSEDE support, in meeting the most challenging scientific problems.
M. Brian Couger, Lenore Pipes, Fabio Squina, Rolf Prade, Adam C. Siepel, Robert Palermo, Michael G. Katze, Christopher E. Mason, Philip D. Blood
Concurr. Comput. Pract. Exp.5
2011 PHAST and RPHAST: phylogenetic analysis with space/time models
abstract
The PHylogenetic Analysis with Space/Time models (PHAST) software package consists of a collection of command-line programs and supporting libraries for comparative genomics. PHAST is best known as the engine behind the Conservation tracks in the University of California, Santa Cruz (UCSC) Genome Browser. However, it also includes several other tools for phylogenetic modeling and functional element identification, as well as utilities for manipulating alignments, trees and genomic annotations. PHAST has been in development since 2002 and has now been downloaded more than 1000 times, but so far it has been released only as provisional ('beta') software. Here, we describe the first official release (v1.0) of PHAST, with improved stability, portability and documentation and several new features. We outline the components of the package and detail recent improvements. In addition, we introduce a new interface to the PHAST libraries from the R statistical computing environment, called RPHAST, and illustrate its use in a series of vignettes. We demonstrate that RPHAST can be particularly useful in applications involving both large-scale phylogenomics and complex statistical analyses. The R interface also makes the PHAST libraries acccessible to non-C programmers, and is useful for rapid prototyping. PHAST v1.0 and RPHAST v1.0 are available for download at http://compgen.bscb.cornell.edu/phast, under the terms of an unrestrictive BSD-style license. RPHAST can also be obtained from the Comprehensive R Archive Network (CRAN; http://cran.r-project.org).
Melissa J. Hubisz, Katherine S. Pollard, Adam C. Siepel
Briefings Bioinform.3
2008 Reconstructing the Evolutionary History of Complex Human Gene Clusters
Yu Zhang 0002, Giltae Song, Tomás Vinar, Eric D. Green, Adam C. Siepel, Webb Miller
RECOMB5
2006 New Methods for Detecting Lineage-Specific Selection
Adam C. Siepel, Katherine S. Pollard, David Haussler
RECOMB1
2006 Identification and Classification of Conserved RNA Secondary Structures in the Human Genome
abstract
The discoveries of microRNAs and riboswitches, among others, have shown functional RNAs to be biologically more important and genomically more prevalent than previously anticipated. We have developed a general comparative genomics method based on phylogenetic stochastic context-free grammars for identifying functional RNAs encoded in the human genome and used it to survey an eight-way genome-wide alignment of the human, chimpanzee, mouse, rat, dog, chicken, zebra-fish, and puffer-fish genomes for deeply conserved functional RNAs. At a loose threshold for acceptance, this search resulted in a set of 48,479 candidate RNA structures. This screen finds a large number of known functional RNAs, including 195 miRNAs, 62 histone 3'UTR stem loops, and various types of known genetic recoding elements. Among the highest-scoring new predictions are 169 new miRNA candidates, as well as new candidate selenocysteine insertion sites, RNA editing hairpins, RNAs involved in transcript auto regulation, and many folds that form singletons or small functional RNA families of completely unknown function. While the rate of false positives in the overall set is difficult to estimate and is likely to be substantial, the results nevertheless provide evidence for many new human functional RNAs and present specific predictions to facilitate their further characterization.
Jakob Skou Pedersen, Gill Bejerano, Adam C. Siepel, Kate R. Rosenbloom, Kerstin Lindblad-Toh, Eric S. Lander, Jim Kent, Webb Miller, David Haussler
PLoS Comput. Biol.3
2004 Computational identification of evolutionarily conserved exons
abstract
Phylogenetic hidden Markov models (phylo-HMMs) have recently been proposed as a means for addressing a multi-species version of the ab initio gene prediction problem. These models allow sequence divergence, a phylogeny, patterns of substitution, and base composition all to be considered simultaneously, in a single unified probabilistic model. Here, we apply phylo-HMMs to a restricted version of the gene prediction problem in which individual exons are sought that are evolutionarily conserved across a diverse set of species. We discuss two new methods for improving prediction performance: (1) the use of context-dependent phylogenetic models, which capture phenomena such as a strong CpG effect in noncoding regions and a preference for synonymous rather than nonsynonymous substitutions in coding regions; and (2) a novel strategy for incorporating insertions and deletion (indels) into the state-transition structure of the model, which captures the different characteristic patterns of alignment gaps in coding and noncoding regions. We also discuss the technique, previously used in pairwise gene predictors, of explicitly modeling conserved noncoding sequence to help reduce false positive predictions. These methods have been incorporated into an exon prediction program called ExoniPhy, and tested with two large data sets. Experimental results indicate that all three methods produce significant improvements in prediction performance. In combination, they lead to prediction accuracy comparable to that of some of the best available gene predictors, despite several limitations of our current models.
Adam C. Siepel, David Haussler
RECOMB1
2003 Combining phylogenetic and hidden Markov models in biosequence analysis
abstract
A few models have appeared in recent years that consider not only the way substitutions occur through evolutionary history at each site of a genome, but also the way the process changes from one site to the next. These models combine phylogenetic models of molecular evolution, which apply to individual sites, and hidden Markov models, which allow for changes from site to site. Besides improving the realism of ordinary phylogenetic models, they are potentially very powerful tools for inference and prediction---for gene finding, for example, or prediction of secondary structure. In this paper, we review progress on combined phylogenetic and hidden Markov models and present some extensions to previous work. Our main result is a simple and efficient method for accommodating higher-order states in the HMM, which allows for context-sensitive models of substitution---that is, models that consider the effects of neighboring bases on the pattern of substitution. We present experimental results indicating that higher-order states, autocorrelated rates, and multiple functional categories all lead to significant improvements in the fit of a combined phylogenetic and hidden Markov model, with the effect of higher-order states being particularly pronounced.
Adam C. Siepel, David Haussler
RECOMB1
2002 An algorithm to enumerate all sorting reversals
abstract
The problem of estimating evolutionary distance from differences in gene order has been distilled to the problem of finding the reversal distance between two signed permutations. During the last decade, much progress was made both in computing reversal distance and in finding a minimum sequence of sorting reversals. For most problem instances, however, many minimum sequences of sorting reversals exist, and obtaining the complete set can be useful in exploring the space of genome rearrangements (e.g., in pursuit of solutions to higher-level problems). The problem of finding all minimum sequences of sorting reversals reduces easily to the problem of finding all sorting reversals of one permutation with respect to another. We derive an efficient algorithm to solve this latter problem, and present experimental results indicating that our algorithm offers a dramatic improvement over the best known alternative. It should be noted that in asymptotic terms the new algorithm does not represent a significant improvement: it requires O(n3) time (where n is the permutation size), while the problem can now be solved trivially in &THgr;(n3) time.
Adam C. Siepel
RECOMB1
2002 Inversion Medians Outperform Breakpoint Medians in Phylogeny Reconstruction from Gene-Order Data
Bernard M. E. Moret, Adam C. Siepel, Jijun Tang
WABI2
2001 Finding an Optimal Inversion Median: Experimental Results
Adam C. Siepel, Bernard M. E. Moret
WABI1
2001 ISYS: a decentralized, component-based approach to the integration of heterogeneous bioinformatics resources
abstract
Abstract Motivation: Heterogeneity of databases and software resources continues to hamper the integration of biological information. Top-down solutions are not feasible for the full-scale problem of integration across biological species and data types. Bottom-up solutions so far have not integrated, in a maximally flexible way, dynamic and interactive graphical-user-interface components with data repositories and analysis tools. Results: We present a component-based approach that relies on a generalized platform for component integration. The platform enables independently-developed components to synchronize their behavior and exchange services, without direct knowledge of one another. An interface-based data model allows the exchange of information with minimal component interdependency. From these interactions an integrated system results, which we call ISYS\batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \(^{TM}\) \end{document}. By allowing services to be discovered dynamically based on selected objects, ISYS encourages a kind of exploratory navigation that we believe to be well-suited for applications in genomic research. Availability: A ‘developer’s kit’ for creating software that uses the ISYS platform is available at www.ncgr.org/research/isys/devrel.html. It includes a more refined and fully documented version of the platform and a starting set of ISYS-ready components. Up-to-date descriptions of the ISYS project can be found at www.ncgr.org/research/isys. Contact: [email protected] * To whom correspondence should be addressed.
Adam C. Siepel, Andrew D. Farmer, Andrew N. Tolopko, Mingzhe Zhuang, Pedro Mendes 0001, William D. Beavis, Bruno W. S. Sobral
Bioinform.1