Jenny C. Taylor

dblp:197/8313 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0003-3602-5704ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
statistical genetics
0.822021
A statistical approach for tracking clonal dynamics in cancer using longitudinal next-generation sequencing data · Bioinform. 2021
Hierarchical probabilistic models for multiple gene/variant associations based on next-generation sequencing data · Bioinform. 2017
Bioinformatics and computational biology
cancer genomics
0.512021
A statistical approach for tracking clonal dynamics in cancer using longitudinal next-generation sequencing data · Bioinform. 2021
Bioinformatics and computational biology › bayesian modeling
dirichlet process mixture model
0.512021
A statistical approach for tracking clonal dynamics in cancer using longitudinal next-generation sequencing data · Bioinform. 2021
Bioinformatics and computational biology › structural bioinformatics
molecular structure visualization
0.412020
MichelaNglo: sculpting protein views on web pages without coding · Bioinform. 2020
Bioinformatics and computational biology
count data modeling
0.312017
Hierarchical probabilistic models for multiple gene/variant associations based on next-generation sequencing data · Bioinform. 2017
Bioinformatics and computational biology › functional genomics
eQTL mapping
0.312017
Hierarchical probabilistic models for multiple gene/variant associations based on next-generation sequencing data · Bioinform. 2017
Bioinformatics and computational biology
genomics
0.312017
ReliableGenome: annotation of genomic regions with high/low variant calling concordance · Bioinform. 2017
Bioinformatics and computational biology › genomics
variant calling
0.312017
ReliableGenome: annotation of genomic regions with high/low variant calling concordance · Bioinform. 2017
Bioinformatics and computational biology › structural bioinformatics
protein structure
0.112020
MichelaNglo: sculpting protein views on web pages without coding · Bioinform. 2020

Methods — techniques the papers use, named apart from their topics

markov chain monte carlo · 0.5gaussian process · 0.5dirichlet process mixture model · 0.5web-based visualization · 0.4sparse bayesian modeling · 0.3laplace smoothing · 0.3consensus calling · 0.3arcsin transformation · 0.3
YearPublicationVenuePosition
2021 A statistical approach for tracking clonal dynamics in cancer using longitudinal next-generation sequencing data
abstract
MOTIVATION: Tumours are composed of distinct cancer cell populations (clones), which continuously adapt to their local micro-environment. Standard methods for clonal deconvolution seek to identify groups of mutations and estimate the prevalence of each group in the tumour, while considering its purity and copy number profile. These methods have been applied on cross-sectional data and on longitudinal data after discarding information on the timing of sample collection. Two key questions are how can we incorporate such information in our analyses and is there any benefit in doing so? RESULTS: We developed a clonal deconvolution method, which incorporates explicitly the temporal spacing of longitudinally sampled tumours. By merging a Dirichlet Process Mixture Model with Gaussian Process priors and using as input a sequence of several sparsely collected samples, our method can reconstruct the temporal profile of the abundance of any mutation cluster supported by the data as a continuous function of time. We benchmarked our method on whole genome, whole exome and targeted sequencing data from patients with chronic lymphocytic leukaemia, on liquid biopsy data from a patient with melanoma and on synthetic data and we found that incorporating information on the timing of tissue collection improves model performance, as long as data of sufficient volume and complexity are available for estimating free model parameters. Thus, our approach is particularly useful when collecting a relatively long sequence of tumour samples is feasible, as in liquid cancers (e.g. leukaemia) and liquid biopsies. AVAILABILITY AND IMPLEMENTATION: The statistical methodology presented in this paper is freely available at github.com/dvav/clonosGP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dimitrios V. Vavoulis, Anthony Cutts, Jenny C. Taylor, Anna Schuh
Bioinform.3
2020 MichelaNglo: sculpting protein views on web pages without coding
abstract
MOTIVATION: The sharing of macromolecular structural information online by scientists is predominantly performed via 2D static images, since the embedding of interactive 3D structures in webpages is non-trivial. Whilst the technologies to do so exist, they are often only implementable with significant web coding experience. RESULTS: Michelaɴɢʟo is an accessible and open-source web-based application that supports the generation, customization and sharing of interactive 3D macromolecular visualizations for digital media without requiring programming skills. A PyMOL file, PDB file, PDB identifier code or protein/gene name can be provided to form the basis of visualizations using the NGL JavaScript library. Hyperlinks that control the view can be added to text within the page. Protein-coding variants can be highlighted to support interpretation of their potential functional consequences. The resulting visualizations and text can be customized and shared, as well as embedded within existing websites by following instructions and using a self-contained download. Michelaɴɢʟo allows researchers to move away from static images and instead engage, describe and explain their protein to a wider audience in a more interactive fashion. AVAILABILITY AND IMPLEMENTATION: Michelaɴɢʟo is hosted at michelanglo.sgc.ox.ac.uk. The Python code is freely available at https://github.com/thesgc/MichelaNGLo, along with documentations about its implementation.
Matteo P. Ferla, Alistair T. Pagnamenta, David R. Damerell, Jenny C. Taylor, Brian D. Marsden
Bioinform.4
2017 ReliableGenome: annotation of genomic regions with high/low variant calling concordance
abstract
MOTIVATION: The increasing adoption of clinical whole-genome resequencing (WGS) demands for highly accurate and reproducible variant calling (VC) methods. The observed discordance between state-of-the-art VC pipelines, however, indicates that the current practice still suffers from non-negligible numbers of false positive and negative SNV and INDEL calls that were shown to be enriched among discordant calls but also in genomic regions with low sequence complexity. RESULTS: Here, we describe our method ReliableGenome (RG) for partitioning genomes into high and low concordance regions with respect to a set of surveyed VC pipelines. Our method combines call sets derived by multiple pipelines from arbitrary numbers of datasets and interpolates expected concordance for genomic regions without data. By applying RG to 219 deep human WGS datasets, we demonstrate that VC concordance depends predominantly on genomic context rather than the actual sequencing data which manifests in high recurrence of regions that can/cannot be reliably genotyped by a single method. This enables the application of pre-computed regions to other data created with comparable sequencing technology and software. RG outperforms comparable efforts in predicting VC concordance and false positive calls in low-concordance regions which underlines its usefulness for variant filtering, annotation and prioritization. RG allows focusing resource-intensive algorithms (e.g. consensus calling methods) on the smaller, discordant share of the genome (20-30%) which might result in increased overall accuracy at reasonable costs. Our method and analysis of discordant calls may further be useful for development, benchmarking and optimization of VC algorithms and for the relative comparison of call sets between different studies/pipelines. AVAILABILITY AND IMPLEMENTATION: RG was implemented in Java, source code and binaries are freely available for non-commercial use at https://github.com/popitsch/wtchg-rg/ CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Niko Popitsch, Anna Schuh, Jenny C. Taylor
Bioinform.4
2017 Hierarchical probabilistic models for multiple gene/variant associations based on next-generation sequencing data
abstract
MOTIVATION: The identification of genetic variants influencing gene expression (known as expression quantitative trait loci or eQTLs) is important in unravelling the genetic basis of complex traits. Detecting multiple eQTLs simultaneously in a population based on paired DNA-seq and RNA-seq assays employs two competing types of models: models which rely on appropriate transformations of RNA-seq data (and are powered by a mature mathematical theory), or count-based models, which represent digital gene expression explicitly, thus rendering such transformations unnecessary. The latter constitutes an immensely popular methodology, which is however plagued by mathematical intractability. RESULTS: We develop tractable count-based models, which are amenable to efficient estimation through the introduction of latent variables and the appropriate application of recent statistical theory in a sparse Bayesian modelling framework. Furthermore, we examine several transformation methods for RNA-seq read counts and we introduce arcsin, logit and Laplace smoothing as preprocessing steps for transformation-based models. Using natural and carefully simulated data from the 1000 Genomes and gEUVADIS projects, we benchmark both approaches under a variety of scenarios, including the presence of noise and violation of basic model assumptions. We demonstrate that an arcsin transformation of Laplace-smoothed data is at least as good as state-of-the-art models, particularly at small samples. Furthermore, we show that an over-dispersed Poisson model is comparable to the celebrated Negative Binomial, but much easier to estimate. These results provide strong support for transformation-based versus count-based (particularly Negative-Binomial-based) models for eQTL mapping. AVAILABILITY AND IMPLEMENTATION: All methods are implemented in the free software eQTLseq: https://github.com/dvav/eQTLseq. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dimitrios V. Vavoulis, Jenny C. Taylor, Anna Schuh
Bioinform.2