VLDB 2026 Research / reviewers in the wild / expert
Taane G. Clark
dblp:45/4405
· DBLP profile ↗
16ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-8985-9265ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Malaria-GENOMAP: a web-based tool for exploring genomic variation of malaria parasitesabstractMOTIVATION: Malaria, caused by Plasmodium parasites, imposes a significant public health burden. While Plasmodium falciparum remains the primary target of elimination strategies due to its high mortality rate, lesser-known species such as P. malariae, P. vivax, and P. knowlesi continue to contribute to substantial human morbidity. Genomic approaches, including whole-genome sequencing, offer powerful tools for understanding the biology, transmission, and emerging drug resistance of these neglected Plasmodium species. However, there is an urgent need for informatic tools to summarize and visualize the high-dimensional and complex genomic data generated. RESULTS: We developed Malaria-GENOMAP, a user-friendly web-based tool, which integrates genomic variant data, such as allele frequencies, with geographical maps and chromosome-wide to gene views for in-depth exploration. The tool includes variation from P. knowlesi (n = 139), P. malariae (n = 158), P. ovale curtisi (n = 36), P. ovale wallikeri (n = 47), P. simium (n = 38), and P. vivax (n = 1359). It enables the investigation of population structure, geographic associations of mutations, and putative drug resistance markers, offering valuable insights for malaria control efforts. AVAILABILITY AND IMPLEMENTATION: Malaria-GENOMAP is available online at https://genomics.lshtm.ac.uk/malaria-genomaps. Joseph Thorpe, Nina Billows, Gabrielle C. Ngwana-Joseph, Amy Ibrahim, Deborah Nolder, Colin J. Sutherland, Thi Hong Ngoc Nguyen, Thi Huong Binh Nguyen, Quang Thieu Nguyen, Jamille G. Dombrowski, Silvia Maria Di Santi, Claudio R. F. Marinho, Jody Phelan, Tomasz J. Kurowski, Fady R. Mohareb, Susana G. Campino, Taane G. Clark |
Bioinform. | 17 |
| 2023 | Feature weighted models to address lineage dependency in drug-resistance prediction from Mycobacterium tuberculosis genome sequencesabstractMOTIVATION: Tuberculosis (TB) is caused by members of the Mycobacterium tuberculosis complex (MTBC), which has a strain- or lineage-based clonal population structure. The evolution of drug-resistance in the MTBC poses a threat to successful treatment and eradication of TB. Machine learning approaches are being increasingly adopted to predict drug-resistance and characterize underlying mutations from whole genome sequences. However, such approaches may not generalize well in clinical practice due to confounding from the population structure of the MTBC. RESULTS: To investigate how population structure affects machine learning prediction, we compared three different approaches to reduce lineage dependency in random forest (RF) models, including stratification, feature selection, and feature weighted models. All RF models achieved moderate-high performance (area under the ROC curve range: 0.60-0.98). First-line drugs had higher performance than second-line drugs, but it varied depending on the lineages in the training dataset. Lineage-specific models generally had higher sensitivity than global models which may be underpinned by strain-specific drug-resistance mutations or sampling effects. The application of feature weights and feature selection approaches reduced lineage dependency in the model and had comparable performance to unweighted RF models. AVAILABILITY AND IMPLEMENTATION: https://github.com/NinaMercedes/RF_lineages. Nina Billows, Jody Phelan, Dong Xia, Yonghong Peng, Taane G. Clark, Yu-Mei Chang |
Bioinform. | 5 |
| 2023 | Bayesian compositional regression with microbiome features via variational inferenceabstractThe microbiome plays a key role in the health of the human body. Interest often lies in finding features of the microbiome, alongside other covariates, which are associated with a phenotype of interest. One important property of microbiome data, which is often overlooked, is its compositionality as it can only provide information about the relative abundance of its constituting components. Typically, these proportions vary by several orders of magnitude in datasets of high dimensions. To address these challenges we develop a Bayesian hierarchical linear log-contrast model which is estimated by mean field Monte-Carlo co-ordinate ascent variational inference (CAVI-MC) and easily scales to high dimensional data. We use novel priors which account for the large differences in scale and constrained parameter space associated with the compositional covariates. A reversible jump Monte Carlo Markov chain guided by the data through univariate approximations of the variational posterior probability of inclusion, with proposal parameters informed by approximating variational densities via auxiliary parameters, is used to estimate intractable marginal expectations. We demonstrate that our proposed Bayesian method performs favourably against existing frequentist state of the art compositional data analysis methods. We then apply the CAVI-MC to the analysis of real data exploring the relationship of the gut microbiome to body mass index. Darren A. V. Scott, Ernest D. Benavente, Julian Libiseller-Egger, Dmitry V. Fedorov, Jody Phelan, Elena Ilina, Polina O. Tikhonova, Alexander Kudryavstev, Julia Galeeva, Taane G. Clark, Alex Lewin |
BMC Bioinform. | 10 |
| 2022 | Portable sequencing of Mycobacterium tuberculosis for clinical and epidemiological applicationsabstractWith >1 million associated deaths in 2020, human tuberculosis (TB) caused by the bacteria Mycobacterium tuberculosis remains one of the deadliest infectious diseases. A plethora of genomic tools and bioinformatics pipelines have become available in recent years to assist the whole genome sequencing of M. tuberculosis. The Oxford Nanopore Technologies (ONT) portable sequencer is a promising platform for cost-effective application in clinics, including personalizing treatment through detection of drug resistance-associated mutations, or in the field, to assist epidemiological and transmission investigations. In this study, we performed a comparison of 10 clinical isolates with DNA sequenced on both long-read ONT and (gold standard) short-read Illumina HiSeq platforms. Our analysis demonstrates the robustness of the ONT variant calling for single nucleotide polymorphisms, despite the high error rate. Moreover, because of improved coverage in repetitive regions where short sequencing reads fail to align accurately, ONT data analysis can incorporate additional regions of the genome usually excluded (e.g. pe/ppe genes). The resulting extra resolution can improve the characterization of transmission clusters and dynamics based on inferring closely related isolates. High concordance in variants in loci associated with drug resistance supports its use for the rapid detection of resistant mutations. Overall, ONT sequencing is a promising tool for TB genomic investigations, particularly to inform clinical and surveillance decision-making to reduce the disease burden. Paula J. Gómez-González, Susana G. Campino, Jody Phelan, Taane G. Clark |
Briefings Bioinform. | 4 |
| 2022 | COVID-profiler: a webserver for the analysis of SARS-CoV-2 sequencing dataabstractBACKGROUND: SARS-CoV-2 virus sequencing has been applied to track the COVID-19 pandemic spread and assist the development of PCR-based diagnostics, serological assays, and vaccines. With sequencing becoming routine globally, bioinformatic tools are needed to assist in the robust processing of resulting genomic data. RESULTS: We developed a web-based bioinformatic pipeline ("COVID-Profiler") that inputs raw or assembled sequencing data, displays raw alignments for quality control, annotates mutations found and performs phylogenetic analysis. The pipeline software can be applied to other (re-) emerging pathogens. CONCLUSIONS: The webserver is available at http://genomics.lshtm.ac.uk/ . The source code is available at https://github.com/jodyphelan/covid-profiler . Jody Phelan, Wouter Deelder, Daniel Ward, Susana G. Campino, Martin L. Hibberd, Taane G. Clark |
BMC Bioinform. | 6 |
| 2020 | Robust detection of point mutations involved in multidrug-resistant Mycobacterium tuberculosis in the presence of co-occurrent resistance markersabstractTuberculosis disease is a major global public health concern and the growing prevalence of drug-resistant Mycobacterium tuberculosis is making disease control more difficult. However, the increasing application of whole-genome sequencing as a diagnostic tool is leading to the profiling of drug resistance to inform clinical practice and treatment decision making. Computational approaches for identifying established and novel resistance-conferring mutations in genomic data include genome-wide association study (GWAS) methodologies, tests for convergent evolution and machine learning techniques. These methods may be confounded by extensive co-occurrent resistance, where statistical models for a drug include unrelated mutations known to be causing resistance to other drugs. Here, we introduce a novel 'cannibalistic' elimination algorithm ("Hungry, Hungry SNPos") that attempts to remove these co-occurrent resistant variants. Using an M. tuberculosis genomic dataset for the virulent Beijing strain-type (n = 3,574) with phenotypic resistance data across five drugs (isoniazid, rifampicin, ethambutol, pyrazinamide, and streptomycin), we demonstrate that this new approach is considerably more robust than traditional methods and detects resistance-associated variants too rare to be likely picked up by correlation-based techniques like GWAS. Julian Libiseller-Egger, Jody Phelan, Susana G. Campino, Fady R. Mohareb, Taane G. Clark |
PLoS Comput. Biol. | 5 |
| 2019 | PrimedRPA: primer design for recombinase polymerase amplification assaysabstractSUMMARY: Recombinase polymerase amplification (RPA), an isothermal nucleic acid amplification method, is enhancing our ability to detect a diverse array of pathogens, thereby assisting the diagnosis of infectious diseases and the detection of microorganisms in food and water. However, new bioinformatics tools are needed to automate and improve the design of the primers and probes sets to be used in RPA, particularly to account for the high genetic diversity of circulating pathogens and cross detection of genetically similar organisms. PrimedRPA is a python-based package that automates the creation and filtering of RPA primers and probe sets. It aligns several sequences to identify conserved targets, and filters regions that cross react with possible background organisms. AVAILABILITY AND IMPLEMENTATION: PrimedRPA was implemented in Python 3 and supported on Linux and MacOS and is freely available from http://pathogenseq.lshtm.ac.uk/PrimedRPA.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matthew Higgins, Matt Ravenhall, Daniel Ward, Jody Phelan, Amy Ibrahim, Matthew S. Forrest, Taane G. Clark, Susana G. Campino |
Bioinform. | 7 |
| 2019 | SV-Pop: population-based structural variant analysis and visualizationabstractBACKGROUND: Genetic structural variation underpins a multitude of phenotypes, with significant implications for a range of biological outcomes. Despite their crucial role, structural variants (SVs) are often neglected and overshadowed by single nucleotide polymorphisms (SNPs), which are used in large-scale analysis such as genome-wide association and population genetic studies. RESULTS: To facilitate the high-throughput analysis of structural variation we have developed an analytical pipeline and visualisation tool, called SV-Pop. The utility of this pipeline was then demonstrated through application with a large, multi-population P. falciparum dataset. CONCLUSIONS: Designed to facilitate downstream analysis and visualisation post-discovery, SV-Pop allows for straightforward integration of multi-population analysis, method and sample-based concordance metrics, and signals of selection. Matt Ravenhall, Susana G. Campino, Taane G. Clark |
BMC Bioinform. | 3 |
| 2015 | PhyTB: Phylogenetic tree visualisation and sample positioning for M. tuberculosisabstractBACKGROUND: Phylogenetic-based classification of M. tuberculosis and other bacterial genomes is a core analysis for studying evolutionary hypotheses, disease outbreaks and transmission events. Whole genome sequencing is providing new insights into the genomic variation underlying intra- and inter-strain diversity, thereby assisting with the classification and molecular barcoding of the bacteria. One roadblock to strain investigation is the lack of user-interactive solutions to interrogate and visualise variation within a phylogenetic tree setting. RESULTS: We have developed a web-based tool called PhyTB ( http://pathogenseq.lshtm.ac.uk/phytblive/index.php ) to assist phylogenetic tree visualisation and identification of M. tuberculosis clade-informative polymorphism. Variant Call Format files can be uploaded to determine a sample position within the tree. A map view summarises the geographical distribution of alleles and strain-types. The utility of the PhyTB is demonstrated on sequence data from 1,601 M. tuberculosis isolates. CONCLUSION: PhyTB contextualises M. tuberculosis genomic variation within epidemiological, geographical and phylogenic settings. Further tool utility is possible by incorporating large variants and phenotypic data (e.g. drug-resistance profiles), and an assessment of genotype-phenotype associations. Source code is available to develop similar websites for other organisms ( http://sourceforge.net/projects/phylotrack ). Ernest D. Benavente, Francesc Coll, Nick Furnham, Ruth McNerney, Judith R. Glynn, Susana G. Campino, Arnab Pain, Fady R. Mohareb, Taane G. Clark |
BMC Bioinform. | 9 |
| 2014 | estMOI: estimating multiplicity of infection using parasite deep sequencing dataabstractAbstract Summary: Individuals living in endemic areas generally harbour multiple parasite strains. Multiplicity of infection (MOI) can be an indicator of immune status and transmission intensity. It has a potentially confounding effect on a number of population genetic analyses, which often assume isolates are clonal. Polymerase chain reaction-based approaches to estimate MOI can lack sensitivity. For example, in the human malaria parasite Plasmodium falciparum, genotyping of the merozoite surface protein (MSP1/2) genes is a standard method for assessing MOI, despite the apparent problem of underestimation. The availability of deep coverage data from massively parallizable sequencing technologies means that MOI can be detected genome wide by considering the abundance of heterozygous genotypes. Here, we present a method to estimate MOI, which considers unique combinations of polymorphisms from sequence reads. The method is implemented within the estMOI software. When applied to clinical P.falciparum isolates from three continents, we find that multiple infections are common, especially in regions with high transmission. Availability and implementation: estMOI is freely available from http://pathogenseq.lshtm.ac.uk. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Samuel A. Assefa, Mark D. Preston, Susana G. Campino, Harold Ocholla, Colin J. Sutherland, Taane G. Clark |
Bioinform. | 6 |
| 2014 | SVAMP: sequence variation analysis, maps and phylogenyabstractSUMMARY: SVAMP is a stand-alone desktop application to visualize genomic variants (in variant call format) in the context of geographical metadata. Users of SVAMP are able to generate phylogenetic trees and perform principal coordinate analysis in real time from variant call format (VCF) and associated metadata files. Allele frequency map, geographical map of isolates, Tajima's D metric, single nucleotide polymorphism density, GC and variation density are also available for visualization in real time. We demonstrate the utility of SVAMP in tracking a methicillin-resistant Staphylococcus aureus outbreak from published next-generation sequencing data across 15 countries. We also demonstrate the scalability and accuracy of our software on 245 Plasmodium falciparum malaria isolates from three continents. AVAILABILITY AND IMPLEMENTATION: The Qt/C++ software code, binaries, user manual and example datasets are available at http://cbrc.kaust.edu.sa/svamp CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Raeece Naeem, Lailatul Hidayah, Mark D. Preston, Taane G. Clark, Arnab Pain |
Bioinform. | 4 |
| 2012 | SpolPred: rapid and accurate prediction of Mycobacterium tuberculosis spoligotypes from short genomic sequencesabstractSUMMARY: Spoligotyping is a well-established genotyping technique based on the presence of unique DNA sequences in Mycobacterium tuberculosis (Mtb), the causal agent of tuberculosis disease (TB). Although advances in sequencing technologies are leading to whole-genome bacterial characterization, tens of thousands of isolates have been spoligotyped, giving a global view of Mtb strain diversity. To bridge the gap, we have developed SpolPred, a software to predict the spoligotype from raw sequence reads. Our approach is compared with experimentally and de novo assembly determined strain types in a set of 44 Mtb isolates. In silico and experimental results are identical for almost all isolates (39/44). However, SpolPred detected five experimentally false spoligotypes and was more accurate and faster than the assembling strategy. Application of SpolPred to an additional seven isolates with no laboratory data led to types that clustered with identical experimental types in a phylogenetic analysis using single-nucleotide polymorphisms. Our results demonstrate the usefulness of the tool and its role in revealing experimental limitations. AVAILABILITY AND IMPLEMENTATION: SpolPred is written in C and is available from www.pathogenseq.org/spolpred. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online. Francesc Coll, Kim Mallard, Mark D. Preston, Stephen D. Bentley, Julian Parkhill, Ruth McNerney, Nigel J. Martin 0001, Taane G. Clark |
Bioinform. | 8 |
| 2012 | VarB: a variation browsing and analysis tool for variants derived from next-generation sequencing dataabstractSUMMARY: There is an immediate need for tools to both analyse and visualize in real-time single-nucleotide polymorphisms, insertions and deletions, and other structural variants from new sequence file formats. We have developed VarB software that can be used to visualize variant call format files in real time, as well as identify regions under balancing selection and informative markers to differentiate user-defined groups (e.g. populations). We demonstrate its utility using sequence data from 50 Plasmodium falciparum isolates comprising two different continents and confirm known signals from genomic regions that contain important antigenic and anti-malarial drug-resistance genes. AVAILABILITY AND IMPLEMENTATION: The C++-based software VarB and user manual are available from www.pathogenseq.org/varb. CONTACT: [email protected] Mark D. Preston, Magnus Manske, Neil R. Horner, Samuel A. Assefa, Susana G. Campino, Sarah Auburn, Issaka Zongo, Jean-Bosco Ouedraogo, Francois Nosten, Timothy J. C. Anderson, Taane G. Clark |
Bioinform. | 11 |
| 2010 | A Bayesian approach using covariance of single nucleotide polymorphism data to detect differences in linkage disequilibrium patterns between groups of individualsabstractMOTIVATION: Quantifying differences in linkage disequilibrium (LD) between sub-groups can highlight genetic regions or sites under selection and/or associated with disease, and may have utility in trans-ethnic mapping studies. RESULTS: We present a novel pseudo Bayes factor (PBF) approach that assess differences in covariance of genotype frequencies from single nucleotide polymorphism (SNP) data from a genome-wide study. The magnitude of the PBF reflects the strength of evidence for a difference, while accounting for the sample size and number of SNPs, without the requirement for permutation testing to establish statistical significance. Application of the PBF to HapMap and Gambian malaria SNP data reveals regional LD differences, some known to be under selection. AVAILABILITY AND IMPLEMENTATION: The PBF approach has been implemented in the BALD (Bayesian analysis of LD differences) C++ software, and is available from http://homepages.lshtm.ac.uk/tgclark/downloads. Taane G. Clark, Susana G. Campino, Elisa Anastasi, Sarah Auburn, Yik Y. Teo, Kerrin S. Small, Kirk A. Rockett, Dominic Kwiatkowski, Christopher C. Holmes |
Bioinform. | 1 |
| 2009 | SnoopCGH: software for visualizing comparative genomic hybridization dataabstractUNLABELLED: Array-based comparative genomic hybridization (CGH) technology is used to discover and validate genomic structural variation, including copy number variants, insertions, deletions and other structural variants (SVs). The visualization and summarization of the array CGH data outputs, potentially across many samples, is an important process in the identification and analysis of SVs. We have developed a software tool for SV analysis using data from array CGH technologies, which is also amenable to short-read sequence data. AVAILABILITY AND IMPLEMENTATION: SnoopCGH is written in java and is available from http://snoopcgh.sourceforge.net/ Jacob Almagro-Garcia, Magnus Manske, Celine Carret, Susana G. Campino, Sarah Auburn, Bronwyn L. MacInnis, Gareth Maslen, Arnab Pain, Christopher I. Newbold, Dominic Kwiatkowski, Taane G. Clark |
Bioinform. | 11 |
| 2007 | A genotype calling algorithm for the Illumina BeadArray platformabstractMOTIVATION: Large-scale genotyping relies on the use of unsupervised automated calling algorithms to assign genotypes to hybridization data. A number of such calling algorithms have been recently established for the Affymetrix GeneChip genotyping technology. Here, we present a fast and accurate genotype calling algorithm for the Illumina BeadArray genotyping platforms. As the technology moves towards assaying millions of genetic polymorphisms simultaneously, there is a need for an integrated and easy-to-use software for calling genotypes. RESULTS: We have introduced a model-based genotype calling algorithm which does not rely on having prior training data or require computationally intensive procedures. The algorithm can assign genotypes to hybridization data from thousands of individuals simultaneously and pools information across multiple individuals to improve the calling. The method can accommodate variations in hybridization intensities which result in dramatic shifts of the position of the genotype clouds by identifying the optimal coordinates to initialize the algorithm. By incorporating the process of perturbation analysis, we can obtain a quality metric measuring the stability of the assigned genotype calls. We show that this quality metric can be used to identify SNPs with low call rates and accuracy. AVAILABILITY: The C++ executable for the algorithm described here is available by request from the authors. Yik Y. Teo, Michael Inouye, Kerrin S. Small, Rhian Gwilliam, Panagiotis Deloukas, Dominic Kwiatkowski, Taane G. Clark |
Bioinform. | 7 |