Tavis K. Anderson

dblp:98/11538 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-3138-5535ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Phylo-rs: an extensible phylogenetic analysis library in rust
abstract
BACKGROUND: The advent of next-generation and long-read sequencing technologies has provided an ever-increasing wealth of phylogenetic data that require specially designed algorithms to decipher the underlying evolutionary relationships. As large-scale data become increasingly accessible, there is a concomitant need for efficient computational libraries that facilitate the development and dissemination of specialized algorithms for phylogenetic comparative biology. RESULTS: We introduce Phylo-rs: a fast, extensible, general-purpose library for phylogenetic analysis and inference written in the Rust programming language. Phylo-rs leverages a combination of speed, memory-safety, and native WebAssembly support offered by Rust to provide a robust set of memory-efficient data structures and elementary phylogenetic algorithms. Phylo-rs focuses on the efficient and convenient deployment of software aimed at large-scale phylogenetic analysis and inference. Scalability analysis against popular libraries shows that Phylo-rs performs comparably or better on key algorithms. We utilized it to assess the phylogenetic diversity of influenza A virus in swine, identifying virus groups that are undergoing evolutionary expansion that could be targeted for control through multivalent vaccines. Additionally, we used Phylo-rs to enhance phylogenetic inference by visualizing tree space from Markov chain Monte Carlo (MCMC) Bayesian analysis, efficiently computing approximately five billion tree pair distances to evaluate convergence and select MCMC runs for genomic epidemiology. CONCLUSION: Phylo-rs enables the design and implementation of cutting-edge software for phylogenetic analysis, thereby facilitating the application and dissemination of theoretical advancements in biology. Phylo-rs is available under an open-source license on GitHub at https://github.com/sriram98v/phylo-rs , with documentation available at https://docs.rs/phylo/latest/phylo/ .
Sriram Vijendran, Tavis K. Anderson, Alexey Markin, Oliver Eulenstein
BMC Bioinform.2
2023 Phylogenetic diversity statistics for all clades in a phylogeny
abstract
The classic quantitative measure of phylogenetic diversity (PD) has been used to address problems in conservation biology, microbial ecology, and evolutionary biology. PD is the minimum total length of the branches in a phylogeny required to cover a specified set of taxa on the phylogeny. A general goal in the application of PD has been identifying a set of taxa of size k that maximize PD on a given phylogeny; this has been mirrored in active research to develop efficient algorithms for the problem. Other descriptive statistics, such as the minimum PD, average PD, and standard deviation of PD, can provide invaluable insight into the distribution of PD across a phylogeny (relative to a fixed value of k). However, there has been limited or no research on computing these statistics, especially when required for each clade in a phylogeny, enabling direct comparisons of PD between clades. We introduce efficient algorithms for computing PD and the associated descriptive statistics for a given phylogeny and each of its clades. In simulation studies, we demonstrate the ability of our algorithms to analyze large-scale phylogenies with applications in ecology and evolutionary biology. The software is available at https://github.com/flu-crew/PD_stats.
Siddhant Grover, Alexey Markin, Tavis K. Anderson, Oliver Eulenstein
Bioinform.3
2022 RF-Net 2: fast inference of virus reassortment and hybridization networks
abstract
MOTIVATION: A phylogenetic network is a powerful model to represent entangled evolutionary histories with both divergent (speciation) and convergent (e.g. hybridization, reassortment, recombination) evolution. The standard approach to inference of hybridization networks is to (i) reconstruct rooted gene trees and (ii) leverage gene tree discordance for network inference. Recently, we introduced a method called RF-Net for accurate inference of virus reassortment and hybridization networks from input gene trees in the presence of errors commonly found in phylogenetic trees. While RF-Net demonstrated the ability to accurately infer networks with up to four reticulations from erroneous input gene trees, its application was limited by the number of reticulations it could handle in a reasonable amount of time. This limitation is particularly restrictive in the inference of the evolutionary history of segmented RNA viruses such as influenza A virus (IAV), where reassortment is one of the major mechanisms shaping the evolution of these pathogens. RESULTS: Here, we expand the functionality of RF-Net that makes it significantly more applicable in practice. Crucially, we introduce a fast extension to RF-Net, called Fast-RF-Net, that can handle large numbers of reticulations without sacrificing accuracy. In addition, we develop automatic stopping criteria to select the appropriate number of reticulations heuristically and implement a feature for RF-Net to output error-corrected input gene trees. We then conduct a comprehensive study of the original method and its novel extensions and confirm their efficacy in practice using extensive simulation and empirical IAV evolutionary analyses. AVAILABILITY AND IMPLEMENTATION: RF-Net 2 is available at https://github.com/flu-crew/rf-net-2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexey Markin, Sanket Wagle, Tavis K. Anderson, Oliver Eulenstein
Bioinform.3
2021 Tripal, a community update after 10 years of supporting open source, standards-based genetic, genomic and breeding databases
abstract
Online, open access databases for biological knowledge serve as central repositories for research communities to store, find and analyze integrated, multi-disciplinary datasets. With increasing volumes, complexity and the need to integrate genomic, transcriptomic, metabolomic, proteomic, phenomic and environmental data, community databases face tremendous challenges in ongoing maintenance, expansion and upgrades. A common infrastructure framework using community standards shared by many databases can reduce development burden, provide interoperability, ensure use of common standards and support long-term sustainability. Tripal is a mature, open source platform built to meet this need. With ongoing improvement since its first release in 2009, Tripal provides full functionality for searching, browsing, loading and curating numerous types of data and is a primary technology powering at least 31 publicly available databases spanning plants, animals and human data, primarily storing genomics, genetics and breeding data. Tripal software development is managed by a shared, inclusive governance structure including both project management and advisory teams. Here, we report on the most important and innovative aspects of Tripal after 11 years development, including integration of diverse types of biological data, successful collaborative projects across member databases, and support for implementing FAIR principles.
Margaret Staton, Ethalinda Cannon, Lacey-Anne Sanderson, Jill L. Wegrzyn, Tavis K. Anderson, Sean Buehler, Irene Cobo-Simón, Kay Faaberg, Emily S. Grau, Valentin Guignon, Jessica Gunoskey, Blake Inderski, Sook Jung, Kelly Lager, Dorrie Main, Monica Poelchau, Risharde Ramnath, Peter Richter 0001, Joe West, Stephen P. Ficklin
Briefings Bioinform.5
2021 Consensus of All Solutions for Intractable Phylogenetic Tree Inference
abstract
Solving median tree problems is a classic approach for inferring species trees from a collection of discordant gene trees. Median tree problems are typically NP-hard and dealt with by local search heuristics. Unfortunately, such heuristics generally lack provable correctness and precision. Algorithmic advances addressing this uncertainty have led to exact dynamic programming formulations suitable to solve a well-studied group of median tree problems for smaller phylogenetic analyses. However, these formulations allow computing only very few optimal species trees out of possibly many such trees, and phylogenetic studies often require the analysis of all optimal solutions through their consensus tree. Here, we describe a significant algorithmic modification of the dynamic programming formulations that compute the cluster counts of all optimal species trees from which various types of consensus trees can be efficiently computed. Through experimental studies, we demonstrate that our parallel implementation of the modified dynamic programming formulation is more efficient than a previous implementation of the original formulation. Finally, we show that the parallel implementation can rapidly identify novel reassorted influenza A viruses potentially facilitating pandemic preparedness efforts.
Pawel Tabaszewski, Pawel Górecki 0001, Alexey Markin, Tavis K. Anderson, Oliver Eulenstein
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 ISU FLUture: a veterinary diagnostic laboratory web-based platform to monitor the temporal genetic patterns of Influenza A virus in swine
abstract
BACKGROUND: Influenza A Virus (IAV) causes respiratory disease in swine and is a zoonotic pathogen. Uncontrolled IAV in swine herds not only affects animal health, it also impacts production through increased costs associated with treatment and prevention efforts. The Iowa State University Veterinary Diagnostic Laboratory (ISU VDL) diagnoses influenza respiratory disease in swine and provides epidemiological analyses on samples submitted by veterinarians. DESCRIPTION: To assess the incidence of IAV in swine and inform stakeholders, the ISU FLUture website was developed as an interactive visualization tool that allows the exploration of the ISU VDL swine IAV aggregate data in the clinical diagnostic database. The information associated with diagnostic cases has varying levels of completeness and is anonymous, but minimally contains: sample collection date, specimen type, and IAV subtype. Many IAV positive samples are sequenced, and in these cases, the hemagglutinin (HA) sequence and genetic classification are completed. These data are collected and presented on ISU FLUture in near real-time, and more than 6,000 IAV positive diagnostic cases and their epidemiological and evolutionary information since 2003 are presented to date. The database and web interface provides rapid and unique insight into the trends of IAV derived from both large- and small-scale swine farms across the United States of America. CONCLUSION: ISU FLUture provides a suite of web-based tools to allow stakeholders to search for trends and correlations in IAV case metadata in swine from the ISU VDL. Since the database infrastructure is updated in near real-time and is integrated within a high-volume veterinary diagnostic laboratory, earlier detection is now possible for emerging IAV in swine that subsequently cause vaccination and control challenges. The access to real-time swine IAV data provides a link with the national USDA swine IAV surveillance system and allows veterinarians to make objective decisions regarding the management and control of IAV in swine. The website is publicly accessible at http://influenza.cvm.iastate.edu .
Michael A. Zeller, Tavis K. Anderson, Rasna W Walia, Amy L. Vincent, Phillip C. Gauger
BMC Bioinform.2
2012 Ranking viruses: measures of positional importance within networks define core viruses for rational polyvalent vaccine development
abstract
MOTIVATION: The extraordinary genetic and antigenic variability of RNA viruses is arguably the greatest challenge to the development of broadly effective vaccines. No single viral variant can induce sufficiently broad immunity, and incorporating all known naturally circulating variants into one multivalent vaccine is not feasible. Furthermore, no objective strategies currently exist to select actual viral variants that should be included or excluded in polyvalent vaccines. RESULTS: To address this problem, we demonstrate a method based on graph theory that quantifies the relative importance of viral variants. We demonstrate our method through application to the envelope glycoprotein gene of a particularly diverse RNA virus of pigs: porcine reproductive and respiratory syndrome virus (PRRSV). Using distance matrices derived from sequence nucleotide difference, amino acid difference and evolutionary distance, we constructed viral networks and used common network statistics to assign each sequence an objective ranking of relative 'importance'. To validate our approach, we use an independent published algorithm to score our top-ranked wild-type variants for coverage of putative T-cell epitopes across the 9383 sequences in our dataset. Top-ranked viruses achieve significantly higher coverage than low-ranked viruses, and top-ranked viruses achieve nearly equal coverage as a synthetic mosaic protein constructed in silico from the same set of 9383 sequences. CONCLUSION: Our approach relies on the network structure of PRRSV but applies to any diverse RNA virus because it identifies subsets of viral variants that are most important to overall viral diversity. We suggest that this method, through the objective quantification of variant importance, provides criteria for choosing viral variants for further characterization, diagnostics, surveillance and ultimately polyvalent vaccine development.
Tavis K. Anderson, William W. Laegreid, Francesco Cerutti, Fernando A. Osorio, Eric A. Nelson, Jane Christopher-Hennings, Tony L. Goldberg
Bioinform.1