EDBT 2026 Demo / reviewers in the wild / expert
Reed A. Cartwright
dblp:08/2478
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2018
0000-0002-0837-9380ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics
variant calling |
0.3 | 1 | 2018 | accuMUlate: a mutation caller designed for mutation accumulation experiments · Bioinform. 2018 |
Bioinformatics and computational biology › statistical genetics
genotype calling |
0.3 | 1 | 2017 | Estimating error models for whole genome sequencing using mixtures of Dirichlet-multinomial distributions · Bioinform. 2017 |
Bioinformatics and computational biology › genomics
genotyping |
0.3 | 1 | 2017 | Estimating error models for whole genome sequencing using mixtures of Dirichlet-multinomial distributions · Bioinform. 2017 |
Bioinformatics and computational biology › genomics › genome sequencing
whole genome sequencing |
0.3 | 1 | 2017 | Estimating error models for whole genome sequencing using mixtures of Dirichlet-multinomial distributions · Bioinform. 2017 |
Bioinformatics and computational biology › sequence alignment
pairwise sequence alignment |
0.1 | 1 | 2007 | Ngila: global pairwise alignments with logarithmic and affine gap costs · Bioinform. 2007 |
Bioinformatics and computational biology
sequence alignment |
0.1 | 1 | 2007 | Ngila: global pairwise alignments with logarithmic and affine gap costs · Bioinform. 2007 |
Bioinformatics and computational biology › phylogenetics
phylogenetic inference |
0.1 | 1 | 2005 | DNA assembly with gaps (Dawg): simulating sequence evolution · Bioinform. 2005 |
Bioinformatics and computational biology
sequence simulation |
0.1 | 1 | 2005 | DNA assembly with gaps (Dawg): simulating sequence evolution · Bioinform. 2005 |
Bioinformatics and computational biology
molecular evolution |
0.0 | 1 | 2005 | DNA assembly with gaps (Dawg): simulating sequence evolution · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
probabilistic modeling · 0.3dirichlet-multinomial mixture model · 0.3log-affine gap cost · 0.1parametric bootstrapping · 0.1general time reversible model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | accuMUlate: a mutation caller designed for mutation accumulation experimentsabstractSummary: Mutation accumulation (MA) is the most widely used method for directly studying the effects of mutation. By sequencing whole genomes from MA lines, researchers can directly study the rate and molecular spectra of spontaneous mutations and use these results to understand how mutation contributes to biological processes. At present there is no software designed specifically for identifying mutations from MA lines. Here we describe accuMUlate, a probabilistic mutation caller that reflects the design of a typical MA experiment while being flexible enough to accommodate properties unique to any particular experiment. Availability and implementation accuMUlate is available from https://github.com/dwinter/accuMUlate. Supplementary information: Supplementary data are available at Bioinformatics online. David J. Winter, Steven H. Wu, Abigail A. Howell, Ricardo B. R. Azevedo, Rebecca A. Zufall, Reed A. Cartwright |
Bioinform. | 6 |
| 2017 | Estimating error models for whole genome sequencing using mixtures of Dirichlet-multinomial distributionsabstractMOTIVATION: Accurate identification of genotypes is an essential part of the analysis of genomic data, including in identification of sequence polymorphisms, linking mutations with disease and determining mutation rates. Biological and technical processes that adversely affect genotyping include copy-number-variation, paralogous sequences, library preparation, sequencing error and reference-mapping biases, among others. RESULTS: We modeled the read depth for all data as a mixture of Dirichlet-multinomial distributions, resulting in significant improvements over previously used models. In most cases the best model was comprised of two distributions. The major-component distribution is similar to a binomial distribution with low error and low reference bias. The minor-component distribution is overdispersed with higher error and reference bias. We also found that sites fitting the minor component are enriched for copy number variants and low complexity regions, which can produce erroneous genotype calls. By removing sites that do not fit the major component, we can improve the accuracy of genotype calls. AVAILABILITY AND IMPLEMENTATION: Methods and data files are available at https://github.com/CartwrightLab/WuEtAl2017/ (doi:10.5281/zenodo.256858). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online. Steven H. Wu, Rachel Schwartz, David J. Winter, Donald F. Conrad, Reed A. Cartwright |
Bioinform. | 5 |
| 2015 | A composite genome approach to identify phylogenetically informative data from next-generation sequencingabstractBACKGROUND: Improvements in sequencing technology now allow easy acquisition of large datasets; however, analyzing these data for phylogenetics can be challenging. We have developed a novel method to rapidly obtain homologous genomic data for phylogenetics directly from next-generation sequencing reads without the use of a reference genome. This software, called SISRS, avoids the time consuming steps of de novo whole genome assembly, multiple genome alignment, and annotation. RESULTS: For simulations SISRS is able to identify large numbers of loci containing variable sites with phylogenetic signal. For genomic data from apes, SISRS identified thousands of variable sites, from which we produced an accurate phylogeny. Finally, we used SISRS to identify phylogenetic markers that we used to estimate the phylogeny of placental mammals. We recovered eight phylogenies that resolved the basal relationships among mammals using datasets with different levels of missing data. The three alternate resolutions of the basal relationships are consistent with the major hypotheses for the relationships among mammals, all of which have been supported previously by different molecular datasets. CONCLUSIONS: SISRS has the potential to transform phylogenetic research. This method eliminates the need for expensive marker development in many studies by using whole genome shotgun sequence data directly. SISRS is open source and freely available at https://github.com/rachelss/SISRS/releases. Rachel Schwartz, Kelly Harkins, Anne Stone, Reed A. Cartwright |
BMC Bioinform. | 4 |
| 2011 | PICS-Ord: Unlimited Coding of Ambiguous Regions by Pairwise Identity and Cost Scores OrdinationabstractBACKGROUND: We present a novel method to encode ambiguously aligned regions in fixed multiple sequence alignments by 'Pairwise Identity and Cost Scores Ordination' (PICS-Ord). The method works via ordination of sequence identity or cost scores matrices by means of Principal Coordinates Analysis (PCoA). After identification of ambiguous regions, the method computes pairwise distances as sequence identities or cost scores, ordinates the resulting distance matrix by means of PCoA, and encodes the principal coordinates as ordered integers. Three biological and 100 simulated datasets were used to assess the performance of the new method. RESULTS: Including ambiguous regions coded by means of PICS-Ord increased topological accuracy, resolution, and bootstrap support in real biological and simulated datasets compared to the alternative of excluding such regions from the analysis a priori. In terms of accuracy, PICS-Ord performs equal to or better than previously available methods of ambiguous region coding (e.g., INAASE), with the advantage of a practically unlimited alignment size and increased analytical speed and the possibility of PICS-Ord scores to be analyzed together with DNA data in a partitioned maximum likelihood model. CONCLUSIONS: Advantages of PICS-Ord over step matrix-based ambiguous region coding with INAASE include a practically unlimited number of OTUs and seamless integration of PICS-Ord codes into phylogenetic datasets, as well as the increased speed of phylogenetic analysis. Contrary to word- and frequency-based methods, PICS-Ord maintains the advantage of pairwise sequence alignment to derive distances, and the method is flexible with respect to the calculation of distance scores. In addition to distance and maximum parsimony, PICS-Ord codes can be analyzed in a Bayesian or maximum likelihood framework. RAxML (version 7.2.6 or higher that was developed for this study) allows up to 32-state ordered or unordered characters. A GTR, MK, or ORDERED model can be applied to analyse the PICS-Ord codes partition, with GTR performing slightly better than MK and ORDERED. AVAILABILITY: An implementation of the PICS-Ord algorithm is available from http://scit.us/projects/ngila/wiki/PICS-Ord. It requires both the statistical software, R http://www.r-project.org and the alignment software Ngila http://scit.us/projects/ngila. Robert K. Luecking, Brendan P. Hodkinson, Alexandros Stamatakis, Reed A. Cartwright |
BMC Bioinform. | 4 |
| 2007 | Ngila: global pairwise alignments with logarithmic and affine gap costsabstractAbstract Summary: Ngila is an application that will find the best alignment of a pair of sequences using log-affine gap costs, which are the most biologically realistic gap costs. Availability: Portable source code for Ngila can be downloaded from its development website, http://scit.us/projects/ngila/. It compiles on most operating systems. Contact: [email protected] or [email protected] Supplementary information: Appendices Reed A. Cartwright |
Bioinform. | 1 |
| 2006 | Logarithmic gap costs decrease alignment accuracyabstractBACKGROUND: Studies on the distribution of indel sizes have consistently found that they obey a power law. This finding has lead several scientists to propose that logarithmic gap costs, G (k) = a + c ln k, are more biologically realistic than affine gap costs, G (k) = a + bk, for sequence alignment. Since quick and efficient affine costs are currently the most popular way to globally align sequences, the goal of this paper is to determine whether logarithmic gap costs improve alignment accuracy significantly enough the merit their use over the faster affine gap costs. RESULTS: A group of simulated sequences pairs were globally aligned using affine, logarithmic, and log-affine gap costs. Alignment accuracy was calculated by comparing resulting alignments to actual alignments of the sequence pairs. Gap costs were then compared based on average alignment accuracy. Log-affine gap costs had the best accuracy, followed closely by affine gap costs, while logarithmic gap costs performed poorly. Subsequently a model was developed to explain the results. CONCLUSION: In contrast to initial expectations, logarithmic gap costs produce poor alignments and are actually not implied by the power-law behavior of gap sizes, given typical match and mismatch costs. Furthermore, affine gap costs not only produce accurate alignments but are also good approximations to biologically realistic gap costs. This work provides added confidence for the biological relevance of existing alignment algorithms. Reed A. Cartwright |
BMC Bioinform. | 1 |
| 2005 | DNA assembly with gaps (Dawg): simulating sequence evolutionabstractMOTIVATION: Relationships amongst taxa are inferred from biological data using phylogenetic methods and procedures. Very few known phylogenies exist against which to test the accuracy of our inferences. Therefore, in the absence of biological data, simulated data must be used to test the accuracy of methods which produce these inferences. Researchers have limited or non-existent options for simulations useful for studying the impact of insertions, deletions, and alignments on phylogenetic accuracy. RESULTS: To satisfy this gap I have developed a new algorithm of indel formation and incorporated it into a new, flexible, and portable application for sequence simulation. The application, called Dawg, simulates phylogenetic evolution of DNA sequences in continuous time using the robust general time reversible model with gamma and invariant rate heterogeneity and a novel length-dependent model of indel formation. On completion, Dawg produces the true alignment of the simulated sequences. Unlike other applications, Dawg allows indel lengths to be explicitly distributed via a biologically realistic power law. Many options are available to allow users to customize their simulations and results. Because simulating with indels would be problematic if biologically realistic parameters could not be estimated, a script is provided with Dawg that can estimate the parameters of indel formation from sequence data. Dawg was applied to the sequences of four chloroplast trnK introns. It was used to parametrically bootstrap an estimation of the rate of indel formation for the phylogeny. Because Dawg can assist in parametric bootstrapping of sequence data it is useful beyond phylogenetics, such as studying alignment algorithms or parameters of molecular evolution. AVAILABILITY: Dawg 1.0.0 can be obtained at the following websites: http://www.genetics.uga.edu/sw/ or http://scit.us/dawg/. The package includes source code, example files, a brief manual and helper scripts. Binary distributions are available for Windows and Macintosh OS X. A development page for Dawg exists at http://scit.us/dawg/, with links to a Subversion repository, mailing lists and updated versions. Reed A. Cartwright |
Bioinform. | 1 |