EDBT 2026 Demo / reviewers in the wild / expert
Christopher Quince
dblp:71/9109
· DBLP profile ↗
7ranked-venue papers
1as first author
1since 2021 · last 2021
0000-0003-1884-8440ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 82% Environmental and earth informatics · 18% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
metagenomics |
0.5 | 3 | 2021 | Accurate Reconstruction of Microbial Strains from Metagenomic Sequencing Using Representative Reference Genomes · RECOMB 2018 Swarm v3: towards tera-scale amplicon clustering · Bioinform. 2021 UCHIME improves sensitivity and speed of chimera detection · Bioinform. 2011 |
Environmental and earth informatics
ecological modeling |
0.3 | 1 | 2017 | Linking Statistical and Ecological Theory: Hubbell's Unified Neutral Theory of Biodiversity as a Hierarchical Dirichlet Process · Proc. IEEE 2017 |
Bioinformatics and computational biology › sequence analysis › sequencing read preprocessing
chimera detection |
0.1 | 1 | 2011 | UCHIME improves sensitivity and speed of chimera detection · Bioinform. 2011 |
Bioinformatics and computational biology
sequence analysis |
0.1 | 1 | 2011 | UCHIME improves sensitivity and speed of chimera detection · Bioinform. 2011 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian computation
bayesian model fitting |
0.1 | 1 | 2017 | Linking Statistical and Ecological Theory: Hubbell's Unified Neutral Theory of Biodiversity as a Hierarchical Dirichlet Process · Proc. IEEE 2017 |
Bioinformatics and computational biology › metagenomics › amplicon sequencing
amplicon sequencing analysis |
0.0 | 1 | 2011 | UCHIME improves sensitivity and speed of chimera detection · Bioinform. 2011 |
Methods — techniques the papers use, named apart from their topics
hierarchical dirichlet process · 0.6bayesian fitting · 0.6multithreading · 0.5c++ optimization · 0.5reference-based chimera detection · 0.1de novo chimera detection · 0.1abundance-based detection · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Swarm v3: towards tera-scale amplicon clusteringabstractMOTIVATION: Previously we presented swarm, an open-source amplicon clustering programme that produces fine-scale molecular operational taxonomic units (OTUs) that are free of arbitrary global clustering thresholds. Here, we present swarm v3 to address issues of contemporary datasets that are growing towards tera-byte sizes. RESULTS: When compared with previous swarm versions, swarm v3 has modernized C++ source code, reduced memory footprint by up to 50%, optimized CPU-usage and multithreading (more than 7 times faster with default parameters), and it has been extensively tested for its robustness and logic. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are available at https://github.com/torognes/swarm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Frédéric Mahé, Lucas Czech, Alexandros Stamatakis, Christopher Quince, Colomban de Vargas, Micah Dunthorn, Torbjørn Rognes |
Bioinform. | 4 |
| 2018 | Accurate Reconstruction of Microbial Strains from Metagenomic Sequencing Using Representative Reference Genomes
Zhemin Zhou, Nina Luhmann, Nabil-Fareed Alikhan, Christopher Quince, Mark Achtman |
RECOMB | 4 |
| 2017 | Linking Statistical and Ecological Theory: Hubbell's Unified Neutral Theory of Biodiversity as a Hierarchical Dirichlet ProcessabstractNeutral models which assume ecological equivalence between species provide null models for community assembly. In Hubbell's unified neutral theory of biodiversity (UNTB), many local communities are connected to a single metacommunity through differing immigration rates. Our ability to fit the full multisite UNTB has hitherto been limited by the lack of a computationally tractable and accurate algorithm. We show that a large class of neutral models with this mainland-island structure but differing local community dynamics converge in the large population limit to the hierarchical Dirichlet process. Using this approximation we developed an efficient Bayesian fitting strategy for the multisite UNTB. We can also use this approach to distinguish between neutral local community assembly given a nonneutral metacommunity distribution and the full UNTB where the metacommunity too assembles neutrally. We applied this fitting strategy to both tropical trees and a data set comprising 570$\,$851 sequences from 278 human gut microbiomes. The tropical tree data set was consistent with the UNTB but for the human gut neutrality was rejected at the whole community level. However, when we applied the algorithm to gut microbial species within the same taxon at different levels of taxonomic resolution, we found that species abundances within some genera were almost consistent with local community assembly. This was not true at higher taxonomic ranks. This suggests that the gut microbiota is more strongly niche constrained than macroscopic organisms, with different groups adopting different functional roles, but within those groups diversity may at least partially be maintained by neutrality. We also observed a negative correlation between body mass index and immigration rates within the family Ruminococcaceae. This provides a novel interpretation of the impact of obesity on the human microbiome as a relative increase in the importance of local growth versus external immigration within this key group of carbohydrate degrading organisms. Keith Harris, Todd L. Parsons, Umer Zeeshan Ijaz, Leo Lahti, Ian H. Holmes, Christopher Quince |
Proc. IEEE | 6 |
| 2016 | Illumina error profiles: resolving fine-scale variation in metagenomic sequencing dataabstractBACKGROUND: Illumina's sequencing platforms are currently the most utilised sequencing systems worldwide. The technology has rapidly evolved over recent years and provides high throughput at low costs with increasing read-lengths and true paired-end reads. However, data from any sequencing technology contains noise and our understanding of the peculiarities and sequencing errors encountered in Illumina data has lagged behind this rapid development. RESULTS: We conducted a systematic investigation of errors and biases in Illumina data based on the largest collection of in vitro metagenomic data sets to date. We evaluated the Genome Analyzer II, HiSeq and MiSeq and tested state-of-the-art low input library preparation methods. Analysing in vitro metagenomic sequencing data allowed us to determine biases directly associated with the actual sequencing process. The position- and nucleotide-specific analysis revealed a substantial bias related to motifs (3mers preceding errors) ending in "GG". On average the top three motifs were linked to 16 % of all substitution errors. Furthermore, a preferential incorporation of ddGTPs was recorded. We hypothesise that all of these biases are related to the engineered polymerase and ddNTPs which are intrinsic to any sequencing-by-synthesis method. We show that quality-score-based error removal strategies can on average remove 69 % of the substitution errors - however, the motif-bias remains. CONCLUSION: Single-nucleotide polymorphism changes in bacterial genomes can cause significant changes in phenotype, including antibiotic resistance and virulence, detecting them within metagenomes is therefore vital. Current error removal techniques are not designed to target the peculiarities encountered in Illumina sequencing data and other sequencing-by-synthesis methods, causing biases to persist and potentially affect any conclusions drawn from the data. In order to develop effective diagnostic and therapeutic approaches we need to be able to identify systematic sequencing errors and distinguish these errors from true genetic variation. Melanie Schirmer, Rosalinda D'Amore, Umer Zeeshan Ijaz, Neil Hall, Christopher Quince |
BMC Bioinform. | 5 |
| 2014 | Benchmarking of viral haplotype reconstruction programmes: an overview of the capacities and limitations of currently available programmesabstractViral haplotype reconstruction from a set of observed reads is one of the most challenging problems in bioinformatics today. Next-generation sequencing technologies enable us to detect single-nucleotide polymorphisms (SNPs) of haplotypes-even if the haplotypes appear at low frequencies. However, there are two major problems. First, we need to distinguish real SNPs from sequencing errors. Second, we need to determine which SNPs occur on the same haplotype, which cannot be inferred from the reads if the distance between SNPs on a haplotype exceeds the read length. We conducted an independent benchmarking study that directly compares the currently available viral haplotype reconstruction programmes. We also present nine in silico data sets that we generated to reflect biologically plausible populations. For these data sets, we simulated 454 and Illumina reads and applied the programmes to test their capacity to reconstruct whole genomes and individual genes. We developed a novel statistical framework to demonstrate the strengths and limitations of the programmes. Our benchmarking demonstrated that all the programmes we tested performed poorly when sequence divergence was low and failed to recover haplotype populations with rare haplotypes. Melanie Schirmer, William T. Sloan, Christopher Quince |
Briefings Bioinform. | 3 |
| 2011 | UCHIME improves sensitivity and speed of chimera detectionabstractMOTIVATION: Chimeric DNA sequences often form during polymerase chain reaction amplification, especially when sequencing single regions (e.g. 16S rRNA or fungal Internal Transcribed Spacer) to assess diversity or compare populations. Undetected chimeras may be misinterpreted as novel species, causing inflated estimates of diversity and spurious inferences of differences between populations. Detection and removal of chimeras is therefore of critical importance in such experiments. RESULTS: We describe UCHIME, a new program that detects chimeric sequences with two or more segments. UCHIME either uses a database of chimera-free sequences or detects chimeras de novo by exploiting abundance data. UCHIME has better sensitivity than ChimeraSlayer (previously the most sensitive database method), especially with short, noisy sequences. In testing on artificial bacterial communities with known composition, UCHIME de novo sensitivity is shown to be comparable to Perseus. UCHIME is >100× faster than Perseus and >1000× faster than ChimeraSlayer. CONTACT: [email protected] AVAILABILITY: Source, binaries and data: http://drive5.com/uchime. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Robert C. Edgar, Brian J. Haas, José Carlos Clemente, Christopher Quince, Rob Knight 0001 |
Bioinform. | 4 |
| 2011 | Removing Noise From Pyrosequenced AmpliconsabstractBACKGROUND: In many environmental genomics applications a homologous region of DNA from a diverse sample is first amplified by PCR and then sequenced. The next generation sequencing technology, 454 pyrosequencing, has allowed much larger read numbers from PCR amplicons than ever before. This has revolutionised the study of microbial diversity as it is now possible to sequence a substantial fraction of the 16S rRNA genes in a community. However, there is a growing realisation that because of the large read numbers and the lack of consensus sequences it is vital to distinguish noise from true sequence diversity in this data. Otherwise this leads to inflated estimates of the number of types or operational taxonomic units (OTUs) present. Three sources of error are important: sequencing error, PCR single base substitutions and PCR chimeras. We present AmpliconNoise, a development of the PyroNoise algorithm that is capable of separately removing 454 sequencing errors and PCR single base errors. We also introduce a novel chimera removal program, Perseus, that exploits the sequence abundances associated with pyrosequencing data. We use data sets where samples of known diversity have been amplified and sequenced to quantify the effect of each of the sources of error on OTU inflation and to validate these algorithms. RESULTS: AmpliconNoise outperforms alternative algorithms substantially reducing per base error rates for both the GS FLX and latest Titanium protocol. All three sources of error lead to inflation of diversity estimates. In particular, chimera formation has a hitherto unrealised importance which varies according to amplification protocol. We show that AmpliconNoise allows accurate estimates of OTU number. Just as importantly AmpliconNoise generates the right OTUs even at low sequence differences. We demonstrate that Perseus has very high sensitivity, able to find 99% of chimeras, which is critical when these are present at high frequencies. CONCLUSIONS: AmpliconNoise followed by Perseus is a very effective pipeline for the removal of noise. In addition the principles behind the algorithms, the inference of true sequences using Expectation-Maximization (EM), and the treatment of chimera detection as a classification or 'supervised learning' problem, will be equally applicable to new sequencing technologies as they appear. Christopher Quince, Anders Lanzén, Russell J. Davenport, Peter J. Turnbaugh |
BMC Bioinform. | 1 |