Rob Knight 0001

dblp:18/2964 · DBLP profile ↗
← Back
31ranked-venue papers
0as first author
8since 2021 · last 2024
0000-0002-0975-9019ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 4 since 2021Artificial intelligence and machine learning · 4Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Mitigating memory latency in FM Index search
abstract
Bowtie 2 is a popular short sequence aligner used by many bioinformatics groups, often as part of the Qiita microbial study management platform. One of the key computational steps during alignment is the extraction of seeds from the read sequences and aligning them with the help of the FM Index. This step was memory access-bound, with access latency as the primary limitation due to a pseudo-random-access pattern. In this paper we present the algorithmic changes used and the performance improvements obtained by re-factoring the code to overlap compute with memory pre-fetching.
Igor Sfiligoi, Daniel McDonald, Rob Knight 0001
e-Science3
2024 Learning the Game: Decoding the Differences between Novice and Expert Players in a Citizen Science Game with Millions of Players
abstract
In recent years, video games have surged in popularity, attracting millions of players across platforms. Citizen science games (CSGs) leverage the processing power of gamers to solve computational and scientific problems. Borderlands Science (BLS) is a mini-game within the mass market game Borderlands 3 that turns multiple sequence alignment (MSA) problems into puzzles. Parallel research demonstrated that BLS players outperformed classical approaches solving small sequence alignment tasks. This study aims to analyze the strategical differences in player solutions in BLS as they gain experience. Through the many collected player solutions from players of different experience level, we gained insights into players’ strategies, differences between expert and non-expert players, and how strategies evolve. We developed a Markov chain trained on solutions from players of different experience levels to understand their actions and outcomes. Results indicate that expert players utilize more gaps and achieve more matches, gradually improving and converging toward unique strategies. Our findings reveal distinct and evolving player strategies. For future citizen science projects, it will be important to consider the identification of player strategies and their evolution over time to improve the game design and data processing.
Eddie Cai, Roman Sarrazin-Gendron, Renata Mutalova, Parham Ghasemloo Gheidari, Alexander Butyaev, Gabriel Richard, Sébastien Caisse, Rob Knight 0001, Mathieu Blanchette, Attila Szantner, Jérôme Waldispühl
FDG8
2024 Scaling DEPP phylogenetic placement to ultra-large reference trees: a tree-aware ensemble approach
abstract
MOTIVATION: Phylogenetic placement of a query sequence on a backbone tree is increasingly used across biomedical sciences to identify the content of a sample from its DNA content. The accuracy of such analyses depends on the density of the backbone tree, making it crucial that placement methods scale to very large trees. Moreover, a new paradigm has been recently proposed to place sequences on the species tree using single-gene data. The goal is to better characterize the samples and to enable combined analyses of marker-gene (e.g., 16S rRNA gene amplicon) and genome-wide data. The recent method DEPP enables performing such analyses using metric learning. However, metric learning is hampered by a need to compute and save a quadratically growing matrix of pairwise distances during training. Thus, the training phase of DEPP does not scale to more than roughly 10 000 backbone species, a problem that we faced when trying to use our recently released Greengenes2 (GG2) reference tree containing 331 270 species. RESULTS: This paper explores divide-and-conquer for training ensembles of DEPP models, culminating in a method called C-DEPP. While divide-and-conquer has been extensively used in phylogenetics, applying divide-and-conquer to data-hungry machine-learning methods needs nuance. C-DEPP uses carefully crafted techniques to enable quasi-linear scaling while maintaining accuracy. C-DEPP enables placing 20 million 16S fragments on the GG2 reference tree in 41 h of computation. AVAILABILITY AND IMPLEMENTATION: The dataset and C-DEPP software are freely available at https://github.com/yueyujiang/dataset_cdepp/.
Yueyu Jiang, Daniel McDonald, Daniela Perry, Rob Knight 0001, Siavash Mirarab
Bioinform.4
2023 Playing the System: Can Puzzle Players Teach us How to Solve Hard Problems?
abstract
With nearly three billion players, video games are more popular than ever. Casual puzzle games are among the most played categories. These games capitalize on the players’ analytical and problem-solving skills. Can we leverage these abilities to teach ourselves how to solve complex combinatorial problems? In this study, we harness the collective wisdom of millions of players to tackle the classical NP-hard problem of multiple sequence alignment, relevant to many areas of biology and medicine. We show that Borderlands Science players propose solutions to multiple sequence alignment tasks that perform as well or better than standard approaches, while exploring a much larger area of the Pareto-optimal solution space. We also show the strategies of the players, although highly heterogeneous, follow a collective logic that can be mimicked with Behavioral Cloning with minimal performance loss, allowing the players’ collective wisdom to be leveraged for alignment of any sequences.
Renata Mutalova, Roman Sarrazin-Gendron, Eddie Cai, Gabriel Richard, Parham Ghasemloo Gheidari, Sébastien Caisse, Rob Knight 0001, Mathieu Blanchette, Attila Szantner, Jérôme Waldispühl
CHI7
2022 SALIENT: Ultra-Fast FPGA-based Short Read Alignment
abstract
State-of-the-art high-throughput DNA sequencers output terabytes of short reads that typically need to be aligned to a reference genome in order to perform downstream analyses. Because alignment typically dominates the total run time of bioinformatics pipelines, a number of recent work sought to accelerate it in hardware. However, existing FPGA implemen-tations did not fully optimize the alignment algorithms for the FPGA hardware and mainly focused on a subset of alignment problems, e.g., ungapped alignment with a limited number of mismatches, which hinder their practical utility. In this work, we analyze the existing alignment methods and identify and leverage opportunities for FPGA acceleration. Our alignment framework, SALIENT, first carries out an ultra-fast ungapped alignment, which supports a flexible number of mismatches. Based on the underlying bioinformatics pipeline and the information provided by the ungapped aligner, SALIENT then identifies a fraction of reads that need to go through its gapped aligner, thus improving alignment throughput. We extensively evaluate SALIENT using diverse datasets. Experimental results indicate that SALIENT, running on a single Xilinx Alveo U280 device, delivers an average throughput of 546 million bases/second, outperforming the state- of-the-art minimap2 software by 40x, and Bowtie2 by up to 107 x, with a similar or slightly better (~O.l %-0.5 %) alignment and error (false negative/positive) rate. Compared to the existing ungapped FPGA aligners [1]–[4], SALIENT has 9.4-18x higher throughput/Watt, while compared to the gapped aligners [5], [6], it is 28–35 x better. SALIENT achieves 7.6 x higher throughput than Illumina DRAGEN Bio-IT Platform [7].
Behnam Khaleghi, Cameron Martino, George Armstrong, Ameen Akel, Ken Curewitz, Justin Eno, Sean Eilert, Rob Knight 0001, Niema Moshiri, Tajana Rosing
FPT9
2021 Galileo: Citizen-led Experimentation Using a Social Computing System
abstract
People have scientific questions and folk theories; yet most lack the expertise to investigate them. How might people transform their questions into experiments that inform both science and their lives? This paper demonstrates how online volunteers can collaboratively design and run experiments using a novel social computing system. The Galileo system provides procedural support using three techniques: 1) experimental design workflow that provides just-in-time training; 2) review workflow with scaffolded questions; and 3) automated routines for data collection. We present two empirical investigations: a study and a field deployment with online volunteers across 16 and 8 countries respectively. People generated structurally-sound experiments on personally meaningful topics; three communities ran a week-long experiment each. We identify two key challenges for citizen-led experimentation—supporting different expertise levels and providing recruitment guidance—and provide specific suggestions from the social computing literature. Our results highlight the promise and challenges of citizen-led knowledge work like experimentation.
Vineet Pandey, Tushar Koul, Daniel McDonald, Madeleine Ball, Bastian Greshake Tzovaras, Rob Knight 0001, Scott R. Klemmer
CHI7
2021 Enabling microbiome research on personal devices
abstract
Microbiome studies have recently transitioned from experimental designs with a few hundred samples to designs spanning tens of thousands of samples. Modern studies such as the Earth Microbiome Project (EMP) afford the statistics crucial for untangling the many factors that influence microbial community composition. Analyzing those data used to require access to a compute cluster, making it both expensive and inconvenient. We show that recent improvements in both hardware and software now allow to compute key bioinformatics tasks on EMP-sized data in minutes using a gaming-class laptop, enabling much faster and broader microbiome science insights.
Igor Sfiligoi, Daniel McDonald, Rob Knight 0001
e-Science3
2021 Experiences and lessons learned from two virtual, hands-on microbiome bioinformatics workshops
abstract
In October of 2020, in response to the Coronavirus Disease 2019 (COVID-19) pandemic, our team hosted our first fully online workshop teaching the QIIME 2 microbiome bioinformatics platform. We had 75 enrolled participants who joined from at least 25 different countries on 6 continents, and we had 22 instructors on 4 continents. In the 5-day workshop, participants worked hands-on with a cloud-based shared compute cluster that we deployed for this course. The event was well received, and participants provided feedback and suggestions in a postworkshop questionnaire. In January of 2021, we followed this workshop with a second fully online workshop, incorporating lessons from the first. Here, we present details on the technology and protocols that we used to run these workshops, focusing on the first workshop and then introducing changes made for the second workshop. We discuss what worked well, what didn't work well, and what we plan to do differently in future workshops.
Matthew R. Dillon, Evan Bolyen, Anja Adamov, Aeriel Belk, Emily Borsom, Zachary Burcham, Justine W. Debelius, Heather Deel, Alex Emmons, Mehrbod Estaki, Chloe Herman, Christopher R. Keefe, Jamie T. Morton, Renato R. M. Oliveira, Andrew Sanchez, Anthony Simard, Yoshiki Vazquez-Baeza, Michal Ziemski, Hazuki E. Miwa, Terry A. Kerere, Carline Coote, Richard Bonneau, Rob Knight 0001, Guilherme C. Oliveira 0001, Piraveen Gopalasingam, Benjamin D. Kaehler, Emily K. Cope, Jessica L. Metcalf, Michael S. Robeson II, Nicholas A. Bokulich, J. Gregory Caporaso
PLoS Comput. Biol.23
2020 SHOGUN: a modular, accurate and scalable framework for microbiome quantification
abstract
SUMMARY: The software pipeline SHOGUN profiles known taxonomic and gene abundances of short-read shotgun metagenomics sequencing data. The pipeline is scalable, modular and flexible. Data analysis and transformation steps can be run individually or together in an automated workflow. Users can easily create new reference databases and can select one of three DNA alignment tools, ranging from ultra-fast low-RAM k-mer-based database search to fully exhaustive gapped DNA alignment, to best fit their analysis needs and computational resources. The pipeline includes an implementation of a published method for taxonomy assignment disambiguation with empirical Bayesian redistribution. The software is installable via the conda resource management framework, has plugins for the QIIME2 and QIITA packages and produces both taxonomy and gene abundance profile tables with a single command, thus promoting convenient and reproducible metagenomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/knights-lab/SHOGUN.
Benjamin Hillmann, Gabriel A. Al-Ghalith, Robin R. Shields-Cutler, Qiyun Zhu, Rob Knight 0001, Dan Knight
Bioinform.5
2019 The genetic basis for adaptation of model-designed syntrophic co-cultures
abstract
Understanding the fundamental characteristics of microbial communities could have far reaching implications for human health and applied biotechnology. Despite this, much is still unknown regarding the genetic basis and evolutionary strategies underlying the formation of viable synthetic communities. By pairing auxotrophic mutants in co-culture, it has been demonstrated that viable nascent E. coli communities can be established where the mutant strains are metabolically coupled. A novel algorithm, OptAux, was constructed to design 61 unique multi-knockout E. coli auxotrophic strains that require significant metabolite uptake to grow. These predicted knockouts included a diverse set of novel non-specific auxotrophs that result from inhibition of major biosynthetic subsystems. Three OptAux predicted non-specific auxotrophic strains-with diverse metabolic deficiencies-were co-cultured with an L-histidine auxotroph and optimized via adaptive laboratory evolution (ALE). Time-course sequencing revealed the genetic changes employed by each strain to achieve higher community growth rates and provided insight into mechanisms for adapting to the syntrophic niche. A community model of metabolism and gene expression was utilized to predict the relative community composition and fundamental characteristics of the evolved communities. This work presents new insight into the genetic strategies underlying viable nascent community formation and a cutting-edge computational method to elucidate metabolic changes that empower the creation of cooperative communities.
Colton J. Lloyd, Zachary A. King, Troy E. Sandberg, Ying Hefner, Connor A. Olson, Patrick V. Phaneuf, Edward J. O'Brien, Jon G. Sanders, Rodolfo A. Salido, Karenina Sanders, Caitriona Brennan, Gregory Humphrey, Rob Knight 0001, Adam M. Feist
PLoS Comput. Biol.13
2019 Ten simple rules for writing and sharing computational analyses in Jupyter Notebooks
abstract
Author(s): Rule, Adam; Birmingham, Amanda; Zuniga, Cristal; Altintas, Ilkay; Huang, Shih-Cheng; Knight, Rob; Moshiri, Niema; Nguyen, Mai H; Rosenthal, Sara Brin; Pérez, Fernando; Rose, Peter W | Editor(s): Lewitter, Fran
Adam Rule, Amanda Birmingham, Cristal Zuñiga, Ilkay Altintas, Shih-Cheng Huang, Rob Knight 0001, Niema Moshiri, Mai H. Nguyen, Sara Brin Rosenthal, Peter W. Rose
PLoS Comput. Biol.6
2018 Docent: transforming personal intuitions to scientific hypotheses through content learning and process training
abstract
People's lived experiences provide intuitions about health. Can they transform these personal intuitions into testable hypotheses that could inform both science and their lives? This paper introduces an online learning architecture and provides system principles for people to brainstorm causal scientific theories. We describe the Learn-Train-Ask workflow that guides participants through learning domain-specific content, process training to frame their intuitions as hypotheses, and collaborating with anonymous peers to brainstorm related questions. 344 voluntary online participants from 27 countries created 399 personally-relevant questions about the human microbiome over 4 months, 75 (19%) of which microbiome experts found potentially scientifically novel. Participants with access to process training generated hypotheses of better quality. Access to learning materials improved the questions' microbiome-specific knowledge. These results highlight the promise of performing personally-meaningful scientific work using massive online learning systems.
Vineet Pandey, Justine W. Debelius, Embriette R. Hyde, Rob Knight 0001, Scott R. Klemmer
L@S5
2017 Gut Instinct: Creating Scientific Theories with Online Learners
abstract
Learners worldwide collectively spend millions of hours per week testing their skills on assignments with known answers. Might some of this time fruitfully be spent posing and exploring novel questions? This paper investigates an approach for learners to contribute scientific ideas. The Gut Instinct system embodies this approach, hosting online learning materials and invites learners to collaboratively brainstorm potential influences on people's microbiome. A between-subjects experiment compared the performance of participants who engaged in just learning, just contributing, or a combination. Participants in the learning condition scored highest on a summative test. Participants in both the contribution and combined conditions generated novel, useful questions; there was not a significant difference between the two. Though participants in the combined condition both learned and contributed, this setting did not exhibit an additive benefit, such as better learning in the combined condition. These results highlight the promise and difficulty of double-bottom-line learning experiences.
Vineet Pandey, Amnon Amir, Justine W. Debelius, Embriette R. Hyde, Rob Knight 0001, Scott R. Klemmer
CHI6
2016 Using machine learning to identify major shifts in human gut microbiome protein family abundance in disease
abstract
Inflammatory Bowel Disease (IBD) is an autoimmune condition that is observed to be associated with major alterations in the gut microbiome taxonomic composition. Here we classify major changes in microbiome protein family abundances between healthy subjects and IBD patients. We use machine learning to analyze results obtained previously from computing relative abundance of ~10,000 KEGG orthologous protein families in the gut microbiome of a set of healthy individuals and IBD patients. We develop a machine learning pipeline, involving the Kolomogorv-Smirnov test, to identify the 100 most statistically significant entries in the KEGG database. Then we use these 100 as a training set for a Random Forest classifier to determine ~5% the KEGGs which are best at separating disease and healthy states. Lastly, we developed a Natural Language Processing classifier of the KEGG description files to predict KEGG relative over-or under-abundance. As we expand our analysis from 10,000 KEGG protein families to one million proteins identified in the gut microbiome, scalable methods for quickly identifying such anomalies between health and disease states will be increasingly valuable for biological interpretation of sequence data.
Mehrdad Yazdani, Bryn C. Taylor, Justine W. Debelius, Rob Knight 0001, Larry Smarr
IEEE BigData5
2015 The unifrac significance test is sensitive to tree topology
abstract
Long et al. (BMC Bioinformatics 2014, 15(1):278) describe a "discrepancy" in using UniFrac to assess statistical significance of community differences. Specifically, they find that weighted UniFrac results differ between input trees where (a) replicate sequences each have their own tip, or (b) all replicates are assigned to one tip with an associated count. We argue that these are two distinct cases that differ in the probability distribution on which the statistical test is based, because of the differences in tree topology. Further study is needed to understand which randomization procedure best detects different aspects of community dissimilarities.
Catherine A. Lozupone, Rob Knight 0001
BMC Bioinform.2
2013 Toward Anthropomimetic Robotics: Development, Simulation, and Control of a Musculoskeletal Torso
abstract
Anthropomimetic robotics differs from conventional approaches by capitalizing on the replication of the inner structures of the human body, such as muscles, tendons, bones, and joints. Here we present our results of more than three years of research in constructing, simulating, and, most importantly, controlling anthropomimetic robots. We manufactured four physical torsos, each more complex than its predecessor, and developed the tools required to simulate their behavior. Furthermore, six different control approaches, inspired by classical control theory, machine learning, and neuroscience, were developed and evaluated via these simulations or in small-scale setups. While the obtained results are encouraging, we are aware that we have barely exploited the potential of the anthropomimetic design so far. But, with the tools developed, we are confident that this novel approach will contribute to our understanding of morphological computation and human motor control in the future.
Steffen Wittmeier, Cristiano Alessandro, Nenad Bascarevic, Konstantinos Dalamagkidis, David Devereux, Alan Diamond, Michael Jäntsch, Kosta Jovanovic, Rob Knight 0001, Hugo Gravato Marques, Predrag Milosavljevic, Bhargav Mitra, Bratislav Svetozarevic, Veljko Potkonjak, Rolf Pfeifer, Alois C. Knoll, Owen Holland
Artif. Life9
2013 A Guide to Enterotypes across the Human Body: Meta-Analysis of Microbial Community Structures in Human Microbiome Datasets
abstract
Recent analyses of human-associated bacterial diversity have categorized individuals into 'enterotypes' or clusters based on the abundances of key bacterial genera in the gut microbiota. There is a lack of consensus, however, on the analytical basis for enterotypes and on the interpretation of these results. We tested how the following factors influenced the detection of enterotypes: clustering methodology, distance metrics, OTU-picking approaches, sequencing depth, data type (whole genome shotgun (WGS) vs.16S rRNA gene sequence data), and 16S rRNA region. We included 16S rRNA gene sequences from the Human Microbiome Project (HMP) and from 16 additional studies and WGS sequences from the HMP and MetaHIT. In most body sites, we observed smooth abundance gradients of key genera without discrete clustering of samples. Some body habitats displayed bimodal (e.g., gut) or multimodal (e.g., vagina) distributions of sample abundances, but not all clustering methods and workflows accurately highlight such clusters. Because identifying enterotypes in datasets depends not only on the structure of the data but is also sensitive to the methods applied to identifying clustering strength, we recommend that multiple approaches be used and compared when testing for enterotypes.
Omry Koren, Dan Knight, Levi Waldron, Nicola Segata, Rob Knight 0001, Curtis Huttenhower, Ruth E. Ley
PLoS Comput. Biol.6
2012 A large-scale benchmark study of existing algorithms for taxonomy-independent microbial community analysis
abstract
Recent advances in massively parallel sequencing technology have created new opportunities to probe the hidden world of microbes. Taxonomy-independent clustering of the 16S rRNA gene is usually the first step in analyzing microbial communities. Dozens of algorithms have been developed in the last decade, but a comprehensive benchmark study is lacking. Here, we survey algorithms currently used by microbiologists, and compare seven representative methods in a large-scale benchmark study that addresses several issues of concern. A new experimental protocol was developed that allows different algorithms to be compared using the same platform, and several criteria were introduced to facilitate a quantitative evaluation of the clustering performance of each algorithm. We found that existing methods vary widely in their outputs, and that inappropriate use of distance levels for taxonomic assignments likely resulted in substantial overestimates of biodiversity in many studies. The benchmark study identified our recently developed ESPRIT-Tree, a fast implementation of the average linkage-based hierarchical clustering algorithm, as one of the best algorithms available in terms of computational efficiency and clustering accuracy.
Yijun Sun, Yunpeng Cai, Susan M. Huse, Rob Knight 0001, William G. Farmerie, Volker Mai
Briefings Bioinform.4
2012 SitePainter: a tool for exploring biogeographical patterns
abstract
UNLABELLED: As microbial ecologists take advantage of high-throughput analytical techniques to describe microbial communities across ever-increasing numbers of samples, the need for new analysis tools that reveal the intrinsic spatial patterns and structures of these populations is crucial. Here we present SitePainter, an interactive graphical tool that allows investigators to create or upload pictures of their study site, load diversity analyses data and display both diversity and taxonomy results in a spatial context. Features of SitePainter include: visualizing α -diversity, using taxonomic summaries; visualizing β -diversity, using results from multidimensional scaling methods; and animating relationships among microbial taxa or pathways overtime. SitePainter thus increases the visual power and ability to explore spatially explicit studies. AVAILABILITY: https://sourceforge.net/projects/sitepainter SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected], [email protected].
Jesse Stombaugh, Christian L. Lauber, Noah Fierer, Rob Knight 0001
Bioinform.5
2011 UCHIME improves sensitivity and speed of chimera detection
abstract
MOTIVATION: Chimeric DNA sequences often form during polymerase chain reaction amplification, especially when sequencing single regions (e.g. 16S rRNA or fungal Internal Transcribed Spacer) to assess diversity or compare populations. Undetected chimeras may be misinterpreted as novel species, causing inflated estimates of diversity and spurious inferences of differences between populations. Detection and removal of chimeras is therefore of critical importance in such experiments. RESULTS: We describe UCHIME, a new program that detects chimeric sequences with two or more segments. UCHIME either uses a database of chimera-free sequences or detects chimeras de novo by exploiting abundance data. UCHIME has better sensitivity than ChimeraSlayer (previously the most sensitive database method), especially with short, noisy sequences. In testing on artificial bacterial communities with known composition, UCHIME de novo sensitivity is shown to be comparable to Perseus. UCHIME is >100× faster than Perseus and >1000× faster than ChimeraSlayer. CONTACT: [email protected] AVAILABILITY: Source, binaries and data: http://drive5.com/uchime. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Robert C. Edgar, Brian J. Haas, José Carlos Clemente, Christopher Quince, Rob Knight 0001
Bioinform.5
2011 TopiaryExplorer: visualizing large phylogenetic trees with environmental metadata
abstract
MOTIVATION: Microbial community profiling is a highly active area of research, but tools that facilitate visualization of phylogenetic trees and associated environmental data have not kept up with the increasing quantity of data generated in these studies. RESULTS: TopiaryExplorer supports the visualization of very large phylogenetic trees, including features such as the automated coloring of branches by environmental data, manipulation of trees and incorporation of per-tip metadata (e.g. taxonomic labels). AVAILABILITY: http://topiaryexplorer.sourceforge.net. CONTACT: [email protected].
Meg Pirrung, Ryan Kennedy, J. Gregory Caporaso, Jesse Stombaugh, Doug Wendel, Rob Knight 0001
Bioinform.6
2011 Boulder ALignment Editor (ALE): a web-based RNA alignment tool
abstract
SUMMARY: The explosion of interest in non-coding RNAs, together with improvements in RNA X-ray crystallography, has led to a rapid increase in RNA structures at atomic resolution from 847 in 2005 to 1900 in 2010. The success of whole-genome sequencing has led to an explosive growth of unaligned homologous sequences. Consequently, there is a compelling and urgent need for user-friendly tools for producing structure-informed RNA alignments. Most alignment software considers the primary sequence alone; some specialized alignment software can also include Watson-Crick base pairs, but none adequately addresses the needs introduced by the rapid influx of both sequence and structural data. Therefore, we have developed the Boulder ALignment Editor (ALE), which is a web-based RNA alignment editor, designed for editing and assessing alignments using structural information. Some features of BoulderALE include the annotation and evaluation of an alignment based on isostericity of Watson-Crick and non-Watson-Crick base pairs, along with the collapsing (horizontally and vertically) of the alignment, while maintaining the ability to edit the alignment. AVAILABILITY: http://www.microbio.me/boulderale.
Jesse Stombaugh, Jeremy Widmann, Daniel McDonald, Rob Knight 0001
Bioinform.4
2011 PrimerProspector: de novo design and taxonomic analysis of barcoded polymerase chain reaction primers
abstract
MOTIVATION: PCR amplification of DNA is a key preliminary step in many applications of high-throughput sequencing technologies, yet design of novel barcoded primers and taxonomic analysis of novel or existing primers remains a challenging task. RESULTS: PrimerProspector is an open-source software package that allows researchers to develop new primers from collections of sequences and to evaluate existing primers in the context of taxonomic data. AVAILABILITY: PrimerProspector is open-source software available at http://pprospector.sourceforge.net CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
William A. Walters, J. Gregory Caporaso, Christian L. Lauber, Donna Berg-Lyons, Noah Fierer, Rob Knight 0001
Bioinform.6
2010 PyNAST: a flexible tool for aligning sequences to a template alignment
abstract
MOTIVATION: The Nearest Alignment Space Termination (NAST) tool is commonly used in sequence-based microbial ecology community analysis, but due to the limited portability of the original implementation, it has not been as widely adopted as possible. Python Nearest Alignment Space Termination (PyNAST) is a complete reimplementation of NAST, which includes three convenient interfaces: a Mac OS X GUI, a command-line interface and a simple application programming interface (API). RESULTS: The availability of PyNAST will make the popular NAST algorithm more portable and thereby applicable to datasets orders of magnitude larger by allowing users to install PyNAST on their own hardware. Additionally because users can align to arbitrary template alignments, a feature not available via the original NAST web interface, the NAST algorithm will be readily applicable to novel tasks outside of microbial community analysis. AVAILABILITY: PyNAST is available at http://pynast.sourceforge.net.
J. Gregory Caporaso, Kyle Bittinger, Frederic D. Bushman, Todd Z. DeSantis, Gary L. Andersen, Rob Knight 0001
Bioinform.6
2009 CodonExplorer: an online tool for analyzing codon usage and sequence composition, scaling from genes to genomes
abstract
DNA composition in general, and codon usage in particular, is crucial for understanding gene function and evolution. CodonExplorer, available online at http://bmf.colorado.edu/codonexplorer/, is an online tool and interactive database that contains millions of genes, allowing rapid exploration of the factors governing gene and genome compositional evolution and exploiting GC content and codon usage frequency to identify genes with composition suggesting high levels of expression or horizontal transfer.
Micah Hamady, Stephanie A. Wilson, Jesse Zaneveld, Noboru Sueoka, Rob Knight 0001
Bioinform.5
2008 Towards mental life as it could be - a robot with imagination
Hugo Gravato Marques, Owen Holland, Rob Knight 0001, Richard A. Newcombe
ALIFE3
2008 Comparison of methods for estimating the nucleotide substitution matrix
abstract
BACKGROUND: The nucleotide substitution rate matrix is a key parameter of molecular evolution. Several methods for inferring this parameter have been proposed, with different mathematical bases. These methods include counting sequence differences and taking the log of the resulting probability matrices, methods based on Markov triples, and maximum likelihood methods that infer the substitution probabilities that lead to the most likely model of evolution. However, the speed and accuracy of these methods has not been compared. RESULTS: Different methods differ in performance by orders of magnitude (ranging from 1 ms to 10 s per matrix), but differences in accuracy of rate matrix reconstruction appear to be relatively small. Encouragingly, relatively simple and fast methods can provide results at least as accurate as far more complex and computationally intensive methods, especially when the sequences to be compared are relatively short. CONCLUSION: Based on the conditions tested, we recommend the use of method of Gojobori et al. (1982) for long sequences (> 600 nucleotides), and the method of Goldman et al. (1996) for shorter sequences (< 600 nucleotides). The method of Barry and Hartigan (1987) can provide somewhat more accuracy, measured as the Euclidean distance between the true and inferred matrices, on long sequences (> 2000 nucleotides) at the expense of substantially longer computation time. The availability of methods that are both fast and accurate will allow us to gain a global picture of change in the nucleotide substitution rate matrix on a genomewide scale across the tree of life.
Maribeth Oscamou, Daniel McDonald, Von Bing Yap, Gavin A. Huttley, Manuel E. Lladser, Rob Knight 0001
BMC Bioinform.6
2008 Pathological rate matrices: from primates to pathogens
abstract
BACKGROUND: Continuous-time Markov models allow flexible, parametrically succinct descriptions of sequence divergence. Non-reversible forms of these models are more biologically realistic but are challenging to develop. The instantaneous rate matrices defined for these models are typically transformed into substitution probability matrices using a matrix exponentiation algorithm that employs eigendecomposition, but this algorithm has characteristic vulnerabilities that lead to significant errors when a rate matrix possesses certain 'pathological' properties. Here we tested whether pathological rate matrices exist in nature, and consider the suitability of different algorithms to their computation. RESULTS: We used concatenated protein coding gene alignments from microbial genomes, primate genomes and independent intron alignments from primate genomes. The Taylor series expansion and eigendecomposition matrix exponentiation algorithms were compared to the less widely employed, but more robust, Padé with scaling and squaring algorithm for nucleotide, dinucleotide, codon and trinucleotide rate matrices. Pathological dinucleotide and trinucleotide matrices were evident in the microbial data set, affecting the eigendecomposition and Taylor algorithms respectively. Even using a conservative estimate of matrix error (occurrence of an invalid probability), both Taylor and eigendecomposition algorithms exhibited substantial error rates: ~100% of all exonic trinucleotide matrices were pathological to the Taylor algorithm while ~10% of codon positions 1 and 2 dinucleotide matrices and intronic trinucleotide matrices, and ~30% of codon matrices were pathological to eigendecomposition. The majority of Taylor algorithm errors derived from occurrence of multiple unobserved states. A small number of negative probabilities were detected from the Padé algorithm on trinucleotide matrices that were attributable to machine precision. Although the Padé algorithm does not facilitate caching of intermediate results, it was up to 3x faster than eigendecomposition on the same matrices. CONCLUSION: Development of robust software for computing non-reversible dinucleotide, codon and higher evolutionary models requires implementation of the Padé with scaling and squaring algorithm.
Harold W. Schranz, Von Bing Yap, Simon Easteal, Rob Knight 0001, Gavin A. Huttley
BMC Bioinform.4
2006 Using the nucleotide substitution rate matrix to detect horizontal gene transfer
abstract
BACKGROUND: Horizontal gene transfer (HGT) has allowed bacteria to evolve many new capabilities. Because transferred genes perform many medically important functions, such as conferring antibiotic resistance, improved detection of horizontally transferred genes from sequence data would be an important advance. Existing sequence-based methods for detecting HGT focus on changes in nucleotide composition or on differences between gene and genome phylogenies; these methods have high error rates. RESULTS: First, we introduce a new class of methods for detecting HGT based on the changes in nucleotide substitution rates that occur when a gene is transferred to a new organism. Our new methods discriminate simulated HGT events with an error rate up to 10 times lower than does GC content. Use of models that are not time-reversible is crucial for detecting HGT. Second, we show that using combinations of multiple predictors of HGT offers substantial improvements over using any single predictor, yielding as much as a factor of 18 improvement in performance (a maximum reduction in error rate from 38% to about 3%). Multiple predictors were combined by using the random forests machine learning algorithm to identify optimal classifiers that separate HGT from non-HGT trees. CONCLUSION: The new class of HGT-detection methods introduced here combines advantages of phylogenetic and compositional HGT-detection techniques. These new techniques offer order-of-magnitude improvements over compositional methods because they are better able to discriminate HGT from non-HGT trees under a wide range of simulated conditions. We also found that combining multiple measures of HGT is essential for detecting a wide range of HGT events. These novel indicators of horizontal transfer will be widely useful in detecting HGT events linked to the evolution of important bacterial traits, such as antibiotic resistance and pathogenicity.
Micah Hamady, Meredith D. Betterton, Rob Knight 0001
BMC Bioinform.3
2006 Fast-Find: A novel computational approach to analyzing combinatorial motifs
abstract
BACKGROUND: Many vital biological processes, including transcription and splicing, require a combination of short, degenerate sequence patterns, or motifs, adjacent to defined sequence features. Although these motifs occur frequently by chance, they only have biological meaning within a specific context. Identifying transcripts that contain meaningful combinations of patterns is thus an important problem, which existing tools address poorly. RESULTS: Here we present a new approach, Fast-FIND (Fast-Fully Indexed Nucleotide Database), that uses a relational database to support rapid indexed searches for arbitrary combinations of patterns defined either by sequence or composition. Fast-FIND is easy to implement, takes less than a second to search the entire Drosophila genome sequence for arbitrary patterns adjacent to sites of alternative polyadenylation, and is sufficiently fast to allow sensitivity analysis on the patterns. We have applied this approach to identify transcripts that contain combinations of sequence motifs for RNA-binding proteins that may regulate alternative polyadenylation. CONCLUSION: Fast-FIND provides an efficient way to identify transcripts that are potentially regulated via alternative polyadenylation. We have used it to generate hypotheses about interactions between specific polyadenylation factors, which we will test experimentally.
Micah Hamady, Erin Peden, Rob Knight 0001, Ravinder Singh
BMC Bioinform.3
2006 UniFrac - An online tool for comparing microbial community diversity in a phylogenetic context
abstract
BACKGROUND: Moving beyond pairwise significance tests to compare many microbial communities simultaneously is critical for understanding large-scale trends in microbial ecology and community assembly. Techniques that allow microbial communities to be compared in a phylogenetic context are rapidly gaining acceptance, but the widespread application of these techniques has been hindered by the difficulty of performing the analyses. RESULTS: We introduce UniFrac, a web application available at http://bmf.colorado.edu/unifrac, that allows several phylogenetic tests for differences among communities to be easily applied and interpreted. We demonstrate the use of UniFrac to cluster multiple environments, and to test which environments are significantly different. We show that analysis of previously published sequences from the Columbia river, its estuary, and the adjacent coastal ocean using the UniFrac interface provided insights that were not apparent from the initial data analysis, which used other commonly employed techniques to compare the communities. CONCLUSION: UniFrac provides easy access to powerful multivariate techniques for comparing microbial communities in a phylogenetic context. We thus expect that it will provide a completely new picture of many microbial interactions and processes in both environmental and medical contexts.
Catherine A. Lozupone, Micah Hamady, Rob Knight 0001
BMC Bioinform.3