Alicia Oshlack

dblp:13/243 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-9788-5690ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 barbieQ: an R software package for analysing barcode count data from clonal tracking experiments
abstract
MOTIVATION: A 'clone' encompasses a progenitor cell and its progeny cells. Tracking clonal composition as cells differentiate or evolve is useful in many fields. Various single-cell lineage tracing (clonal tracking) technologies use unique DNA barcodes that are passed from progenitor cells to their offspring. The barcode count for each sample indicates cell number in clones. However, analysis of barcode count data is often bespoke and relies on visualisations and heuristics. A generalized workflow for preprocessing and robust statistical analysis of barcode count data across protocols is needed. RESULTS: We introduce barbieQ, a Bioconductor R package for analysing barcode count data across groups of samples. It provides data-driven quality control and filtering, extensive visualisations, and two statistical tests: (1) Differential barcode proportion (differences in proportions between sample groups), and (2) Differential barcode occurrence (differences in presence/absence odds between groups). Both tests handle complex experimental designs using regression models and rigorously account for sample-to-sample variability. We validated both tests on semi-simulated, real data and a case study, demonstrating that they hold their size, are sufficiently powered to detect true differences, and outperform existing approaches. AVAILABILITY: barbieQ is available on Bioconductor at https://doi.org/10.18129/B9.bioc.barbieQ.
Liyang Fei, Jovana Maksimovic, Alicia Oshlack
Bioinform.3
2024 Damsel: analysis and visualisation of DamID sequencing in R
abstract
SUMMARY: DamID sequencing is a technique to map the genome-wide interaction of a protein with DNA. Damsel is the first Bioconductor package to provide an end to end analysis for DamID sequencing data within R. Damsel performs quantification and testing of significant binding sites along with exploratory and visual analysis. Damsel produces results consistent with previous analysis approaches. AVAILABILITY AND IMPLEMENTATION: The R package Damsel is available for install through the Bioconductor project https://bioconductor.org/packages/release/bioc/html/Damsel.html and the code is available on GitHub https://github.com/Oshlack/Damsel/.
Caitlin G. Page, Andrew Lonsdale, Katrina A. Mitchell, Jan Schröder, Kieran F. Harvey, Alicia Oshlack
Bioinform.6
2022 propeller: testing for differences in cell type proportions in single cell data
abstract
MOTIVATION: Single cell RNA-Sequencing (scRNA-seq) has rapidly gained popularity over the last few years for profiling the transcriptomes of thousands to millions of single cells. This technology is now being used to analyse experiments with complex designs including biological replication. One question that can be asked from single cell experiments, which has been difficult to directly address with bulk RNA-seq data, is whether the cell type proportions are different between two or more experimental conditions. As well as gene expression changes, the relative depletion or enrichment of a particular cell type can be the functional consequence of disease or treatment. However, cell type proportion estimates from scRNA-seq data are variable and statistical methods that can correctly account for different sources of variability are needed to confidently identify statistically significant shifts in cell type composition between experimental conditions. RESULTS: We have developed propeller, a robust and flexible method that leverages biological replication to find statistically significant differences in cell type proportions between groups. Using simulated cell type proportions data, we show that propeller performs well under a variety of scenarios. We applied propeller to test for significant changes in cell type proportions related to human heart development, ageing and COVID-19 disease severity. AVAILABILITY AND IMPLEMENTATION: The propeller method is publicly available in the open source speckle R package (https://github.com/phipsonlab/speckle). All the analysis code for the article is available at the associated analysis website: https://phipsonlab.github.io/propeller-paper-analysis/. The speckle package, analysis scripts and datasets have been deposited at https://doi.org/10.5281/zenodo.7009042. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Belinda Phipson, Choon Boon Sim, Enzo R. Porrello, Alex W. Hewitt, Joseph E. Powell, Alicia Oshlack
Bioinform.6
2021 Detecting copy number alterations in RNA-Seq using SuperFreq
abstract
MOTIVATION: Calling copy number alterations (CNAs) from RNA sequencing (RNA-Seq) is challenging, because of the marked variability in coverage across genes and paucity of single nucleotide polymorphisms (SNPs). We have adapted SuperFreq to call absolute and allele sensitive CNAs from RNA-Seq. SuperFreq uses an error-propagation framework to combine and maximize information from read counts and B-allele frequencies. RESULTS: We used datasets from The Cancer Genome Atlas (TCGA) to assess the validity of CNA calls from RNA-Seq. When ploidy estimates were consistent, we found agreement with DNA SNP-arrays for over 98% of the genome for acute myeloid leukaemia (TCGA-AML, n = 116) and 87% for colorectal cancer (TCGA-CRC, n = 377). The sensitivity of CNA calling from RNA-Seq was dependent on gene density. Using RNA-Seq, SuperFreq detected 78% of CNA calls covering 100 or more genes with a precision of 94%. Recall dropped for focal events, but this also depended on signal intensity. For example, in the CRC cohort SuperFreq identified all cases (7/7) with high-level amplification of ERBB2, where the copy number was typically >20, but identified only 6% of cases (1/17) with moderate amplification of IGF2, which occurs over a smaller interval. SuperFreq offers an integrated platform for identification of CNAs and point mutations. As evidence of how SuperFreq can be applied, we used it to reproduce the established relationship between somatic mutation load and CNA profile in CRC using RNA-Seq alone. AVAILABILITY AND IMPLEMENTATION: SuperFreq is implemented in R and the code is available through GitHub: https://github.com/ChristofferFlensburg/SuperFreq/. Data and code to reproduce the figures are available at: https://gitlab.wehi.edu.au/flensburg.c/SuperFreq_RNA_paper. Data from TCGA (phs000178) was accessed from GDC following completion of a data access request through the database of Genotypes and Phenotypes (dbGaP). Data from the Leucegene consortium was downloaded from GEO (AML samples: GSE67040; normal CD34+ cells: GSE48846). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Christoffer Flensburg, Alicia Oshlack, Ian J. Majewski
Bioinform.2
2020 SuperFreq: Integrated mutation detection and clonal tracking in cancer
abstract
Analysing multiple cancer samples from an individual patient can provide insight into the way the disease evolves.Monitoring the expansion and contraction of distinct clones helps to reveal the mutations that initiate the disease and those that drive progression.Existing approaches for clonal tracking from sequencing data typically require the user to combine multiple tools that are not purpose-built for this task.Furthermore, most methods require a matched normal (non-tumour) sample, which limits the scope of application.We developed SuperFreq, a cancer exome sequencing analysis pipeline that integrates identification of somatic single nucleotide variants (SNVs) and copy number alterations (CNAs) and clonal tracking for both.SuperFreq does not require a matched normal and instead relies on unrelated controls.When analysing multiple samples from a single patient, SuperFreq cross checks variant calls to improve clonal tracking, which helps to separate somatic from germline variants, and to resolve overlapping CNA calls.To demonstrate our software we analysed 304 cancer-normal exome samples across 33 cancer types in The Cancer Genome Atlas (TCGA) and evaluated the quality of the SNV and CNA calls.We simulated clonal evolution through in silico mixing of cancer and normal samples in known proportion.We found that SuperFreq identified 93% of clones with a cellular fraction of at least 50% and mutations were assigned to the correct clone with high recall and precision.In addition, SuperFreq maintained a similar level of performance for most aspects of the analysis when run without a matched normal.SuperFreq is highly versatile and can be applied in many different experimental settings for the analysis of exomes and other capture libraries.We demonstrate an application of SuperFreq to leukaemia patients with diagnosis and relapse samples. Author summaryCancer is a disease that continues to evolve.Understanding how it changes can provide key biological insights; for example, it can help to identify recurrent patterns associated with therapy resistance.However, tracking clonal evolution in a cancer from sequencing data is a major analytical challenge.We have developed SuperFreq, an analysis framework purpose built for the detection of intra-tumoural heterogeneity and clonal evolution.
Christoffer Flensburg, Tobias Sargeant, Alicia Oshlack, Ian J. Majewski
PLoS Comput. Biol.3
2018 Exploring the single-cell RNA-seq analysis landscape with the scRNA-tools database
abstract
As single-cell RNA-sequencing (scRNA-seq) datasets have become more widespread the number of tools designed to analyse these data has dramatically increased. Navigating the vast sea of tools now available is becoming increasingly challenging for researchers. In order to better facilitate selection of appropriate analysis tools we have created the scRNA-tools database (www.scRNA-tools.org) to catalogue and curate analysis tools as they become available. Our database collects a range of information on each scRNA-seq analysis tool and categorises them according to the analysis tasks they perform. Exploration of this database gives insights into the areas of rapid development of analysis methods for scRNA-seq data. We see that many tools perform tasks specific to scRNA-seq analysis, particularly clustering and ordering of cells. We also find that the scRNA-seq community embraces an open-source and open-science approach, with most tools available under open-source licenses and preprints being extensively used as a means to describe methods. The scRNA-tools database provides a valuable resource for researchers embarking on scRNA-seq analysis and records the growth of the field over time.
Luke Zappia, Belinda Phipson, Alicia Oshlack
PLoS Comput. Biol.3
2016 missMethyl: an R package for analyzing data from Illumina's HumanMethylation450 platform
abstract
UNLABELLED: DNA methylation is one of the most commonly studied epigenetic modifications due to its role in both disease and development. The Illumina HumanMethylation450 BeadChip is a cost-effective way to profile >450 000 CpGs across the human genome, making it a popular platform for profiling DNA methylation. Here we introduce missMethyl, an R package with a suite of tools for performing normalization, removal of unwanted variation in differential methylation analysis, differential variability testing and gene set analysis for the 450K array. AVAILABILITY AND IMPLEMENTATION: missMethyl is an R package available from the Bioconductor project at www.bioconductor.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Belinda Phipson, Jovana Maksimovic, Alicia Oshlack
Bioinform.3
2015 Genotyping microsatellites in next-generation sequencing data
abstract
Microsatellites are short (2-6bp) DNA sequences repeated in tandem, which make up approximately 3% of the human genome [ 1 ]. These loci are prone to frequent mutations and high polymorphism with the estimated mutation rates of 10 - 10 events per locus per generation, orders of magnitude higher than other parts of the genome [ 2 ]. Dozens of neurological and developmental disorders have been attributed to microsatellite expansions [ 3 ]. Microsatellites have also been implicated in a range of functions such as DNA replication and repair, chromatin organisation and regulation of gene expression [ 4 ]. Traditionally, microsatellite variation has been measured using capillary gel electrophoresis [ 5 ]. In addition to being time-consuming, and expensive, this method fails to reveal the full complexity at these loci because it does not directly sequence the fragment but only measure the number of bases in the repeat. Next-generation sequencing has the potential to address these problems. However, determining microsatellite lengths using next-generation sequencing data is difficult. In particular, polymerase slippage during PCR amplification introduces stutter noise. A small number of software tools have been written to genotype simple microsatellites in next-generation sequencing data [ 6 – 8 ], however they fail to address the issues of SNPs and compound repeats, and in some cases provide only approximate genotypes. We have begun to develop a microsatellite genotyping algorithm that addresses these issues, providing high accuracy as well as more detailed analysis of microsatellite loci. We have validated it using high depth amplicon sequencing data of microsatellites near the AVPR1A gene. We found high concordance between our algorithm and repeat lengths obtained by electrophoresis, manual inspection and Mendelian inheritance (Table 1 ). By subsampling the reads, we found that our model is accurate to within one repeat unit down to coverages that we would expect in standard exome sequencing (Figure 1 ). Additionally, we detected polymorphic single nucleotide changes within some microsatellites. Genotyping accuracy at the (AC) n promoter locus as a function of the number of reads spanning the microsatellite . 20 to 3000 reads were sampled with replacement from those spanning the microsatellite. This was done 1000 times for each depth. A shows the portion of genotypes that were exactly correct, B shows the proportion of genotypes that were correct to within one repeat unit. The algorithm was approximately 95% correct at calling the exact same genotype on high depth sequencing data. When it did call a genotype incorrectly, the genotype was only one repeat unit different. The algorithm can perform at approximately 90% accuracy to within one repeat unit with as few as 20 informative reads and reaches almost 100% accuracy to within one repeat unit with 100 or more informative reads. Future work will include expanding the algorithm to genotype compound microsatellites and further validation and comparison with other algorithms will be performed on whole genome data sets.
Harriet Dashnow, Susan Tan, Debjani Das, Simon Easteal, Alicia Oshlack
BMC Bioinform.5
2012 Bpipe: a tool for running and managing bioinformatics pipelines
abstract
SUMMARY: Bpipe is a simple, dedicated programming language for defining and executing bioinformatics pipelines. It specializes in enabling users to turn existing pipelines based on shell scripts or command line tools into highly flexible, adaptable and maintainable workflows with a minimum of effort. Bpipe ensures that pipelines execute in a controlled and repeatable fashion and keeps audit trails and logs to ensure that experimental results are reproducible. Requiring only Java as a dependency, Bpipe is fully self-contained and cross-platform, making it very easy to adopt and deploy into existing environments. AVAILABILITY AND IMPLEMENTATION: Bpipe is freely available from http://bpipe.org under a BSD License.
Simon P. Sadedin, Bernard J. Pope, Alicia Oshlack
Bioinform.3
2007 Using DNA microarrays to study gene expression in closely related species
abstract
MOTIVATION: Comparisons of gene expression levels within and between species have become a central tool in the study of the genetic basis for phenotypic variation, as well as in the study of the evolution of gene regulation. DNA microarrays are a key technology that enables these studies. Currently, however, microarrays are only available for a small number of species. Thus, in order to study gene expression levels in species for which microarrays are not available, researchers face three sets of choices: (i) use a microarray designed for another species, but only compare gene expression levels within species, (ii) construct a new microarray for every species whose gene expression profiles will be compared or (iii) build a multi-species microarray with probes from each species of interest. Here, we use data collected using a multi-primate cDNA array to evaluate the reliability of each approach. RESULTS: We find that, for inter-species comparisons, estimates of expression differences based on multi-species microarrays are more accurate than those based on multiple species-specific arrays. We also demonstrate that within-species expression differences can be estimated using a microarray for a closely related species, without discernible loss of information. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alicia Oshlack, Adrien E. Chabot, Gordon K. Smyth, Yoav Gilad
Bioinform.1
2007 A comparison of background correction methods for two-colour microarrays
abstract
MOTIVATION: Microarray data must be background corrected to remove the effects of non-specific binding or spatial heterogeneity across the array, but this practice typically causes other problems such as negative corrected intensities and high variability of low intensity log-ratios. Different estimators of background, and various model-based processing methods, are compared in this study in search of the best option for differential expression analyses of small microarray experiments. RESULTS: Using data where some independent truth in gene expression is known, eight different background correction alternatives are compared, in terms of precision and bias of the resulting gene expression measures, and in terms of their ability to detect differentially expressed genes as judged by two popular algorithms, SAM and limma eBayes. A new background processing method (normexp) is introduced which is based on a convolution model. The model-based correction methods are shown to be markedly superior to the usual practice of subtracting local background estimates. Methods which stabilize the variances of the log-ratios along the intensity range perform the best. The normexp+offset method is found to give the lowest false discovery rate overall, followed by morph and vsn. Like vsn, normexp is applicable to most types of two-colour microarray data. AVAILABILITY: The background correction methods compared in this article are available in the R package limma (Smyth, 2005) from http://www.bioconductor.org. SUPPLEMENTARY INFORMATION: Supplementary data are available from http://bioinf.wehi.edu.au/resources/webReferences.html.
Matthew E. Ritchie, Jeremy David Silver, Alicia Oshlack, Melissa Holmes, Dileepa S. Diyagama, Andrew J. Holloway, Gordon K. Smyth
Bioinform.3
2006 Statistical analysis of an RNA titration series evaluates microarray precision and sensitivity on a whole-array basis
abstract
BACKGROUND: Concerns are often raised about the accuracy of microarray technologies and the degree of cross-platform agreement, but there are yet no methods which can unambiguously evaluate precision and sensitivity for these technologies on a whole-array basis. RESULTS: A methodology is described for evaluating the precision and sensitivity of whole-genome gene expression technologies such as microarrays. The method consists of an easy-to-construct titration series of RNA samples and an associated statistical analysis using non-linear regression. The method evaluates the precision and responsiveness of each microarray platform on a whole-array basis, i.e., using all the probes, without the need to match probes across platforms. An experiment is conducted to assess and compare four widely used microarray platforms. All four platforms are shown to have satisfactory precision but the commercial platforms are superior for resolving differential expression for genes at lower expression levels. The effective precision of the two-color platforms is improved by allowing for probe-specific dye-effects in the statistical model. The methodology is used to compare three data extraction algorithms for the Affymetrix platforms, demonstrating poor performance for the commonly used proprietary algorithm relative to the other algorithms. For probes which can be matched across platforms, the cross-platform variability is decomposed into within-platform and between-platform components, showing that platform disagreement is almost entirely systematic rather than due to measurement variability. CONCLUSION: The results demonstrate good precision and sensitivity for all the platforms, but highlight the need for improved probe annotation. They quantify the extent to which cross-platform measures can be expected to be less accurate than within-platform comparisons for predicting disease progression or outcome.
Andrew J. Holloway, Alicia Oshlack, Dileepa S. Diyagama, David D. L. Bowtell, Gordon K. Smyth
BMC Bioinform.2