Natalie P. Thorne

dblp:23/3802 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
array CGH analysis
0.112006
BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006
Bioinformatics and computational biology › genomics › computational genomics
copy number segmentation
0.112006
BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006
Bioinformatics and computational biology
genomics
0.112006
BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.012006
BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006

Methods — techniques the papers use, named apart from their topics

heterogeneous hidden markov model · 0.1
YearPublicationVenuePosition
2018 Evaluation of computational programs to predict HLA genotypes from genomic sequencing data
abstract
Motivation: Despite being essential for numerous clinical and research applications, high-resolution human leukocyte antigen (HLA) typing remains challenging and laboratory tests are also time-consuming and labour intensive. With next-generation sequencing data becoming widely accessible, on-demand in silico HLA typing offers an economical and efficient alternative. Results: In this study we evaluate the HLA typing accuracy and efficiency of five computational HLA typing methods by comparing their predictions against a curated set of > 1000 published polymerase chain reaction-derived HLA genotypes on three different data sets (whole genome sequencing, whole exome sequencing and transcriptomic sequencing data). The highest accuracy at clinically relevant resolution (four digits) we observe is 81% on RNAseq data by PHLAT and 99% accuracy by OptiType when limited to Class I genes only. We also observed variability between the tools for resource consumption, with runtime ranging from an average of 5 h (HLAminer) to 7 min (seq2HLA) and memory from 12.8 GB (HLA-VBSeq) to 0.46 GB (HLAminer) per sample. While a minimal coverage is required, other factors also determine prediction accuracy and the results between tools do not correlate well. Therefore, by combining tools, there is the potential to develop a highly accurate ensemble method that is able to deliver fast, economical HLA typing from existing sequencing data.
Denis C. Bauer, Armella Zadoorian, Laurence O. W. Wilson, Melbourne Genomics Health Alliance, Natalie P. Thorne
Briefings Bioinform.5
2007 Missing channels in two-colour microarray experiments: Combining single-channel and two-channel data
abstract
BACKGROUND: There are mechanisms, notably ozone degradation, that can damage a single channel of two-channel microarray experiments. Resulting analyses therefore often choose between the unacceptable inclusion of poor quality data or the unpalatable exclusion of some (possibly a lot of) good quality data along with the bad. Two such approaches would be a single channel analysis using some of the data from all of the arrays, and an analysis of all of the data, but only from unaffected arrays. In this paper we examine a 'combined' approach to the analysis of such affected experiments that uses all of the unaffected data. RESULTS: A simulation experiment shows that while a single channel analysis performs relatively well when the majority of arrays are affected, and excluding affected arrays performs relatively well when few arrays are affected (as would be expected in both cases), the combined approach out-performs both. There are benefits to actively estimating the key-parameter of the approach, but whether these compensate for the increased computational cost and complexity over just setting that parameter to take a fixed value is not clear. Inclusion of ozone-affected data results in poor performance, with a clear spatial effect in the damage being apparent. CONCLUSION: There is no need to exclude unaffected data in order to remove those which are damaged. The combined approach discussed here is shown to out-perform more usual approaches, although it seems that if the damage is limited to very few arrays, or extends to very nearly all, then the benefits will be limited. In other circumstances though, large improvements in performance can be achieved by adopting such an approach.
Andy G. Lynch, David E. Neal, John D. Kelly, Glyn J. Burtt, Natalie P. Thorne
BMC Bioinform.5
2006 BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data
abstract
Abstract Summary: We have developed a new method (BioHMM) for segmenting array comparative genomic hybridization data into states with the same underlying copy number. By utilizing a heterogeneous hidden Markov model, BioHMM incorporates relevant biological factors (e.g. the distance between adjacent clones) in the segmentation process. Availability: BioHMM is available as part of the R library snapCGH which can be downloaded from Contact: [email protected] Supplementary information: Supplementary information is available at
John C. Marioni, Natalie P. Thorne, Simon Tavaré
Bioinform.2