VLDB 2026 Research / reviewers in the wild / expert
Anthony T. Papenfuss
dblp:144/6544
· DBLP profile ↗
9ranked-venue papers
0as first author
5since 2021 · last 2023
0000-0002-1102-8506ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A systematic review of computational methods for designing efficient guides for CRISPR DNA base editor systemsabstractIn only a few years, as a breakthrough technology, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein (CRISPR/Cas) gene-editing systems have ushered in the era of genome engineering with a plethora of applications. One of the most promising CRISPR tools, so-called base editors, opened an exciting avenue for exploring new therapeutic approaches through controlled mutagenesis. However, the efficiency of a base editor guide varies depending on several biological determinants, such as chromatin accessibility, DNA repair proteins, transcriptional activity, factors related to local sequence context and so on. Thus, the success of genetic perturbation directed by CRISPR/Cas base-editing systems relies on an optimal single guide RNA (sgRNA) design, taking those determinants into account. Although there is 11 commonly used software to design guides specifically for base editors, only three of them investigated and implemented those biological determinants into their models. This review presents the key features, capabilities and limitations of all currently available software with a particular focus on predictive model-based algorithms. Here, we summarize existing software for sgRNA design and provide a base for improving the efficiency of existing available software suites for precise target base editing. Göknur Giner, Saima Ikram, Marco J. Herold, Anthony T. Papenfuss |
Briefings Bioinform. | 4 |
| 2023 | cellsig plug-in enhances CIBERSORTx signature selection for multidataset transcriptomes with sparse multilevel modellingabstractMOTIVATION: The precise characterization of cell-type transcriptomes is pivotal to understanding cellular lineages, deconvolution of bulk transcriptomes, and clinical applications. Single-cell RNA sequencing resources like the Human Cell Atlas have revolutionised cell-type profiling. However, challenges persist due to data heterogeneity and discrepancies across different studies. One limitation of prevailing tools such as CIBERSORTx is their inability to address hierarchical data structures and handle nonoverlapping gene sets across samples, relying on filtering or imputation. RESULTS: Here, we present cellsig, a Bayesian sparse multilevel model designed to improve signature estimation by adjusting data for multilevel effects and modelling for gene-set sparsity. Our model is tailored to large-scale, heterogeneous pseudobulk and bulk RNA sequencing data collections with nonoverlapping gene sets. We tested the performances of cellsig on a novel curated Human Bulk Cell-type Catalogue, which harmonizes 1435 samples across 58 datasets. We show that cellsig significantly enhances cell-type marker gene ranking performance. This approach is valuable for cell-type signature selection, with implications for marker gene validation, single-cell annotation, and deconvolution benchmarks. AVAILABILITY AND IMPLEMENTATION: Codes and the interactive app are available at https://github.com/stemangiola/cellsig; and the database is available at https://doi.org/10.5281/zenodo.7582421. Md. Abdullah-Al-Kamran Khan, Yuhan Sun 0001, Alexander D. Barrow, Anthony T. Papenfuss, Stefano Mangiola |
Bioinform. | 5 |
| 2022 | StructuralVariantAnnotation: a R/Bioconductor foundation for a caller-agnostic structural variant software ecosystemabstractSUMMARY: StructuralVariantAnnotation is an R/Bioconductor package that provides a framework for decoupling downstream analysis of structural variant breakpoints from upstream variant calling methods. It standardizes the representational format from BEDPE, or any of the three different notations supported by VCF into a breakpoint GRanges data structure suitable for use by the wider Bioconductor ecosystem. It handles both transitive breakpoints and duplication/insertion notational differences of identical variants-both common scenarios when comparing short/long read-based call sets that confound downstream analysis. StructuralVariantAnnotation provides the caller-agnostic foundation needed for a R/Bioconductor ecosystem of structural variant annotation, classification and interpretation tools able to handle both simple and complex genomic rearrangements. AVAILABILITY AND IMPLEMENTATION: StructuralVariantAnnotation is implemented in R and available for download as the Bioconductor StructuralVariantAnnotation package. Details can be found at https://www.bioconductor.org/packages/release/bioc/html/StructuralVariantAnnotation.html. It has been released under a GPL license. Daniel Cameron, Ruining Dong, Anthony T. Papenfuss |
Bioinform. | 3 |
| 2021 | VIRUSBreakend: Viral Integration Recognition Using Single BreakendsabstractMOTIVATION: Integration of viruses into infected host cell DNA can cause DNA damage and disrupt genes. Recent cost reductions and growth of whole genome sequencing has produced a wealth of data in which viral presence and integration detection is possible. While key research and clinically relevant insights can be uncovered, existing software has not achieved widespread adoption, limited in part due to high computational costs, the inability to detect a wide range of viruses, as well as precision and sensitivity. RESULTS: Here, we describe VIRUSBreakend, a high-speed tool that identifies viral DNA presence and genomic integration. It utilizes single breakends, breakpoints in which only one side can be unambiguously placed, in a novel virus-centric variant calling and assembly approach to identify viral integrations with high sensitivity and a near-zero false discovery rate. VIRUSBreakend detects viral integrations anywhere in the host genome including regions such as centromeres and telomeres unable to be called by existing tools. Applying VIRUSBreakend to a large metastatic cancer cohort, we demonstrate that it can reliably detect clinically relevant viral presence and integration including HPV, HBV, MCPyV, EBV and HHV-8. AVAILABILITY AND IMPLEMENTATION: VIRUSBreakend is part of the Genomic Rearrangement IDentification Software Suite (GRIDSS). It is available under a GPLv3 license from https://github.com/PapenfussLab/VIRUSBreakend. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniel Cameron, Nina Jacobs, Paul Roepman, Peter Priestley, Edwin Cuppen, Anthony T. Papenfuss |
Bioinform. | 6 |
| 2021 | Interfacing Seurat with the R tidy universeabstractMOTIVATION: Seurat is one of the most popular software suites for the analysis of single-cell RNA sequencing data. Considering the popularity of the tidyverse ecosystem, which offers a large set of data display, query, manipulation, integration and visualization utilities, a great opportunity exists to interface the Seurat object with the tidyverse. This interface gives the large data science community of tidyverse users the possibility to operate with familiar grammar. RESULTS: To provide Seurat with a tidyverse-oriented interface without compromising efficiency, we developed tidyseurat, a lightweight adapter to the tidyverse. Tidyseurat displays cell information as a tibble abstraction, allowing intuitively interfacing Seurat with dplyr, tidyr, ggplot2 and plotly packages powering efficient data manipulation, integration and visualization. Iterative analyses on data subsets are enabled by interfacing with the popular nest-map framework. AVAILABILITY AND IMPLEMENTATION: The software is freely available at cran.r-project.org/web/packages/tidyseurat and github.com/stemangiola/tidyseurat. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Stefano Mangiola, Maria A. Doyle, Anthony T. Papenfuss |
Bioinform. | 3 |
| 2017 | HYSYS: have you swapped your samples?abstractMotivation: The application of a genomics assay to samples from a cohort is a frequently applied experimental design in cancer genomics studies. The collection and analysis of cancer sequencing data in the clinical setting is an elaborate process that may involve consenting patients, obtaining possibly-multiple DNA samples, sequencing and analysis. Many of these steps are manual. At any stage mistakes can occur that cause a DNA sample to be labelled incorrectly. However, there is a paucity of methods in the literature to identify such swaps specifically in cancer studies. Results: Here, we introduce a simple method, HYSYS, to estimate the relatedness of samples and test for sample swaps and contamination. The test uses the concordance of homozygous SNPs between samples. The method is motivated by the observation that homozygous germline population variants rarely change in the disease and are not affected by loss of heterozygosity. Our tools include visualization and a testing framework to flag possible sample swaps. We demonstrate the utility of this approach on a small cohort. Availability and Implementation: http://github.com/PapenfussLab/HaveYouSwappedYourSamples. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Jan Schröder, Vincent Corbin, Anthony T. Papenfuss |
Bioinform. | 3 |
| 2017 | Canary: an atomic pipeline for clinical amplicon assaysabstractBACKGROUND: High throughput sequencing requires bioinformatics pipelines to process large volumes of data into meaningful variants that can be translated into a clinical report. These pipelines often suffer from a number of shortcomings: they lack robustness and have many components written in multiple languages, each with a variety of resource requirements. Pipeline components must be linked together with a workflow system to achieve the processing of FASTQ files through to a VCF file of variants. Crafting these pipelines requires considerable bioinformatics and IT skills beyond the reach of many clinical laboratories. RESULTS: Here we present Canary, a single program that can be run on a laptop, which takes FASTQ files from amplicon assays through to an annotated VCF file ready for clinical analysis. Canary can be installed and run with a single command using Docker containerization or run as a single JAR file on a wide range of platforms. Although it is a single utility, Canary performs all the functions present in more complex and unwieldy pipelines. All variants identified by Canary are 3' shifted and represented in their most parsimonious form to provide a consistent nomenclature, irrespective of sequencing variation. Further, proximate in-phase variants are represented as a single HGVS 'delins' variant. This allows for correct nomenclature and consequences to be ascribed to complex multi-nucleotide polymorphisms (MNPs), which are otherwise difficult to represent and interpret. Variants can also be annotated with hundreds of attributes sourced from MyVariant.info to give up to date details on pathogenicity, population statistics and in-silico predictors. CONCLUSIONS: Canary has been used at the Peter MacCallum Cancer Centre in Melbourne for the last 2 years for the processing of clinical sequencing data. By encapsulating clinical features in a single, easily installed executable, Canary makes sequencing more accessible to all pathology laboratories. Canary is available for download as source or a Docker image at https://github.com/PapenfussLab/Canary under a GPL-3.0 License. Kenneth D. Doig, Jason Ellul, Andrew Fellowes, Ella R. Thompson, Georgina L. Ryland, Piers Blombery, Anthony T. Papenfuss, Stephen B. Fox |
BMC Bioinform. | 7 |
| 2017 | CLOVE: classification of genomic fusions into structural variation eventsabstractBACKGROUND: A precise understanding of structural variants (SVs) in DNA is important in the study of cancer and population diversity. Many methods have been designed to identify SVs from DNA sequencing data. However, the problem remains challenging because existing approaches suffer from low sensitivity, precision, and positional accuracy. Furthermore, many existing tools only identify breakpoints, and so not collect related breakpoints and classify them as a particular type of SV. Due to the rapidly increasing usage of high throughput sequencing technologies in this area, there is an urgent need for algorithms that can accurately classify complex genomic rearrangements (involving more than one breakpoint or fusion). RESULTS: We present CLOVE, an algorithm for integrating the results of multiple breakpoint or SV callers and classifying the results as a particular SV. CLOVE is based on a graph data structure that is created from the breakpoint information. The algorithm looks for patterns in the graph that are characteristic of more complex rearrangement types. CLOVE is able to integrate the results of multiple callers, producing a consensus call. CONCLUSIONS: We demonstrate using simulated and real data that re-classified SV calls produced by CLOVE improve on the raw call set of existing SV algorithms, particularly in terms of accuracy. CLOVE is freely available from http://www.github.com/PapenfussLab . Jan Schröder, Adrianto Wirawan, Bertil Schmidt, Anthony T. Papenfuss |
BMC Bioinform. | 4 |
| 2014 | Socrates: identification of genomic rearrangements in tumour genomes by re-aligning soft clipped readsabstractMOTIVATION: Methods for detecting somatic genome rearrangements in tumours using next-generation sequencing are vital in cancer genomics. Available algorithms use one or more sources of evidence, such as read depth, paired-end reads or split reads to predict structural variants. However, the problem remains challenging due to the significant computational burden and high false-positive or false-negative rates. RESULTS: In this article, we present Socrates (SOft Clip re-alignment To idEntify Structural variants), a highly efficient and effective method for detecting genomic rearrangements in tumours that uses only split-read data. Socrates has single-nucleotide resolution, identifies micro-homologies and untemplated sequence at break points, has high sensitivity and high specificity and takes advantage of parallelism for efficient use of resources. We demonstrate using simulated and real data that Socrates performs well compared with a number of existing structural variant detection tools. AVAILABILITY AND IMPLEMENTATION: Socrates is released as open source and available from http://bioinf.wehi.edu.au/socrates CONTACT: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Jan Schröder, Arthur L. Hsu, Samantha E. Boyle, Geoff MacIntyre, Marek Cmero, Richard W. Tothill, Ricky W. Johnstone, Mark Shackleton, Anthony T. Papenfuss |
Bioinform. | 9 |