EDBT 2026 Demo / reviewers in the wild / expert
John C. Marioni
dblp:64/7592
· DBLP profile ↗
8ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0001-9092-0852ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
genomics |
0.9 | 3 | 2023 | RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023 Tools for mapping high-throughput sequencing data · Bioinform. 2012 BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006 |
Bioinformatics and computational biology › genomics › genomic data management
genomic data sharing |
0.7 | 1 | 2023 | RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023 |
Bioinformatics and computational biology
gene expression analysis |
0.2 | 1 | 2016 | HDTD: analyzing multi-tissue gene expression data · Bioinform. 2016 |
Bioinformatics and computational biology › sequence analysis
read mapping |
0.2 | 2 | 2012 | Tools for mapping high-throughput sequencing data · Bioinform. 2012 Effect of read-mapping biases on detecting allele-specific expression from RNA-sequencing data · Bioinform. 2009 |
Bioinformatics and computational biology › genomics
high-throughput sequencing |
0.1 | 1 | 2012 | Tools for mapping high-throughput sequencing data · Bioinform. 2012 |
Bioinformatics and computational biology › gene expression analysis › gene expression quantification
allele-specific expression |
0.1 | 1 | 2009 | Effect of read-mapping biases on detecting allele-specific expression from RNA-sequencing data · Bioinform. 2009 |
Bioinformatics and computational biology › transcriptomics › transcriptome sequencing
RNA sequencing analysis |
0.1 | 1 | 2009 | Effect of read-mapping biases on detecting allele-specific expression from RNA-sequencing data · Bioinform. 2009 |
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics |
0.1 | 2 | 2016 | HDTD: analyzing multi-tissue gene expression data · Bioinform. 2016 BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006 |
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
array CGH analysis |
0.1 | 1 | 2006 | BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006 |
Bioinformatics and computational biology › genomics › computational genomics
copy number segmentation |
0.1 | 1 | 2006 | BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data · Bioinform. 2006 |
Methods — techniques the papers use, named apart from their topics
matrix slicing · 0.7robust estimation · 0.2hypothesis testing · 0.2survey · 0.1simulation · 0.1read mapping · 0.1heterogeneous hidden markov model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | RNAget: an API to securely retrieve RNA quantificationsabstractSUMMARY: Large-scale sharing of genomic quantification data requires standardized access interfaces. In this Global Alliance for Genomics and Health project, we developed RNAget, an API for secure access to genomic quantification data in matrix form. RNAget provides for slicing matrices to extract desired subsets of data and is applicable to all expression matrix-format data, including RNA sequencing and microarrays. Further, it generalizes to quantification matrices of other sequence-based genomics such as ATAC-seq and ChIP-seq. AVAILABILITY AND IMPLEMENTATION: https://ga4gh-rnaseq.github.io/schema/docs/index.html. Sean Upchurch, Emilio Palumbo, Jeremy Adams, David Bujold, Guillaume Bourque, Jared Nedzel, Keenan Graham, Meenakshi S. Kagda, Pedro Assis, Benjamin C. Hitz, Emilio Righi, Roderic Guigó, Barbara J. Wold, Alvis Brazma, Julia Burchard, Joe Capka, Michael Cherry, Laura Clarke, Brian Craft, Manolis Dermitzakis, Mark Diekhans, John Dursi, Michael Sean Fitzsimons, Zac Flaming, Romina Garrido, Alfred Gil, Paul Godden, Matt Green, Mitch Guttman, Brian Haas, Max Haeussler, Sten Linnarsson, Adam Lipski, Simonne Longerich, David R. Lougheed, Jonathan Manning, John C. Marioni, Christopher Meyer, Stephen B. Montgomery, Alyssa Morrow, Alfonso Muñoz-Pomer Fuentes, Jared L. Nedzel, Kevin Osborn, Francis Ouellette, Irene Papatheodorou, Dmitri D. Pervouchine, Arun K. Ramani, Jordi Rambla De Argila, Bashir Sadjad, David Steinberg, Jeremiah Talkar, Timothy Tickle, Kathy Tzeng, Saman Vaisipour, Sean Watford, Barbara Wold |
Bioinform. | 39 |
| 2016 | HDTD: analyzing multi-tissue gene expression dataabstractMOTIVATION: By collecting multiple samples per subject, researchers can characterize intra-subject variation using physiologically relevant measurements such as gene expression profiling. This can yield important insights into fundamental biological questions ranging from cell type identity to tumour development. For each subject, the data measurements can be written as a matrix with the different subsamples (e.g. multiple tissues) indexing the columns and the genes indexing the rows. In this context, neither the genes nor the tissues are expected to be independent and straightforward application of traditional statistical methods that ignore this two-way dependence might lead to erroneous conclusions. Herein, we present a suite of tools embedded within the R/Bioconductor package HDTD for robustly estimating and performing hypothesis tests about the mean relationship and the covariance structure within the rows and columns. We illustrate the utility of HDTD by applying it to analyze data generated by the Genotype-Tissue Expression consortium. AVAILABILITY AND IMPLEMENTATION: The R package HDTD is part of Bioconductor. The source code and a comprehensive user's guide are available at http://bioconductor.org/packages/release/bioc/html/HDTD.html CONTACT: : [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anestis Touloumis, John C. Marioni, Simon Tavaré |
Bioinform. | 2 |
| 2015 | BASiCS: Bayesian Analysis of Single-Cell Sequencing DataabstractSingle-cell mRNA sequencing can uncover novel cell-to-cell heterogeneity in gene expression levels in seemingly homogeneous populations of cells. However, these experiments are prone to high levels of unexplained technical noise, creating new challenges for identifying genes that show genuine heterogeneous expression within the population of cells under study. BASiCS (Bayesian Analysis of Single-Cell Sequencing data) is an integrated Bayesian hierarchical model where: (i) cell-specific normalisation constants are estimated as part of the model parameters, (ii) technical variability is quantified based on spike-in genes that are artificially introduced to each analysed cell's lysate and (iii) the total variability of the expression counts is decomposed into technical and biological components. BASiCS also provides an intuitive detection criterion for highly (or lowly) variable genes within the population of cells under study. This is formalised by means of tail posterior probabilities associated to high (or low) biological cell-to-cell variance contributions, quantities that can be easily interpreted by users. We demonstrate our method using gene expression measurements from mouse Embryonic Stem Cells. Cross-validation and meaningful enrichment of gene ontology categories within genes classified as highly (or lowly) variable supports the efficacy of our approach. Catalina A. Vallejos, John C. Marioni, Sylvia Richardson |
PLoS Comput. Biol. | 2 |
| 2014 | Identifying Cell Types from Spatially Referenced Single-Cell Expression DatasetsabstractComplex tissues, such as the brain, are composed of multiple different cell types, each of which have distinct and important roles, for example in neural function. Moreover, it has recently been appreciated that the cells that make up these sub-cell types themselves harbour significant cell-to-cell heterogeneity, in particular at the level of gene expression. The ability to study this heterogeneity has been revolutionised by advances in experimental technology, such as Wholemount in Situ Hybridizations (WiSH) and single-cell RNA-sequencing. Consequently, it is now possible to study gene expression levels in thousands of cells from the same tissue type. After generating such data one of the key goals is to cluster the cells into groups that correspond to both known and putatively novel cell types. Whilst many clustering algorithms exist, they are typically unable to incorporate information about the spatial dependence between cells within the tissue under study. When such information exists it provides important insights that should be directly included in the clustering scheme. To this end we have developed a clustering method that uses a Hidden Markov Random Field (HMRF) model to exploit both quantitative measures of expression and spatial information. To accurately reflect the underlying biology, we extend current HMRF approaches by allowing the degree of spatial coherency to differ between clusters. We demonstrate the utility of our method using simulated data before applying it to cluster single cell gene expression data generated by applying WiSH to study expression patterns in the brain of the marine annelid Platynereis dumereilii. Our approach allows known cell types to be identified as well as revealing new, previously unexplored cell types within the brain of this important model system. Jean-Baptiste Pettit, Raju Tomer, Kaia Achim, Sylvia Richardson, Lamiae Azizi, John C. Marioni |
PLoS Comput. Biol. | 6 |
| 2013 | bioWeb3D: an online webGL 3D data visualisation toolabstractBACKGROUND: Data visualization is critical for interpreting biological data. However, in practice it can prove to be a bottleneck for non trained researchers; this is especially true for three dimensional (3D) data representation. Whilst existing software can provide all necessary functionalities to represent and manipulate biological 3D datasets, very few are easily accessible (browser based), cross platform and accessible to non-expert users. RESULTS: An online HTML5/WebGL based 3D visualisation tool has been developed to allow biologists to quickly and easily view interactive and customizable three dimensional representations of their data along with multiple layers of information. Using the WebGL library Three.js written in Javascript, bioWeb3D allows the simultaneous visualisation of multiple large datasets inputted via a simple JSON, XML or CSV file, which can be read and analysed locally thanks to HTML5 capabilities. CONCLUSIONS: Using basic 3D representation techniques in a technologically innovative context, we provide a program that is not intended to compete with professional 3D representation software, but that instead enables a quick and intuitive representation of reasonably large 3D datasets. Jean-Baptiste Pettit, John C. Marioni |
BMC Bioinform. | 2 |
| 2012 | Tools for mapping high-throughput sequencing dataabstractAbstract Motivation: A ubiquitous and fundamental step in high-throughput sequencing analysis is the alignment (mapping) of the generated reads to a reference sequence. To accomplish this task, numerous software tools have been proposed. Determining the mappers that are most suitable for a specific application is not trivial. Results: This survey focuses on classifying mappers through a wide number of characteristics. The goal is to allow practitioners to compare the mappers more easily and find those that are most suitable for their specific problem. Availability: A regularly updated compendium of mappers can be found at http://wwwdev.ebi.ac.uk/fg/hts_mappers/. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Nuno A. Fonseca, Johan Rung, Alvis Brazma, John C. Marioni |
Bioinform. | 4 |
| 2009 | Effect of read-mapping biases on detecting allele-specific expression from RNA-sequencing dataabstractMOTIVATION: Next-generation sequencing has become an important tool for genome-wide quantification of DNA and RNA. However, a major technical hurdle lies in the need to map short sequence reads back to their correct locations in a reference genome. Here, we investigate the impact of SNP variation on the reliability of read-mapping in the context of detecting allele-specific expression (ASE). RESULTS: We generated 16 million 35 bp reads from mRNA of each of two HapMap Yoruba individuals. When we mapped these reads to the human genome we found that, at heterozygous SNPs, there was a significant bias toward higher mapping rates of the allele in the reference sequence, compared with the alternative allele. Masking known SNP positions in the genome sequence eliminated the reference bias but, surprisingly, did not lead to more reliable results overall. We find that even after masking, approximately 5-10% of SNPs still have an inherent bias toward more effective mapping of one allele. Filtering out inherently biased SNPs removes 40% of the top signals of ASE. The remaining SNPs showing ASE are enriched in genes previously known to harbor cis-regulatory variation or known to show uniparental imprinting. Our results have implications for a variety of applications involving detection of alternate alleles from short-read sequence data. AVAILABILITY: Scripts, written in Perl and R, for simulating short reads, masking SNP variation in a reference genome and analyzing the simulation output are available upon request from JFD. Raw short read data were deposited in GEO (http://www.ncbi.nlm.nih.gov/geo/) under accession number GSE18156. CONTACT: [email protected]; [email protected]; [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jacob F. Degner, John C. Marioni, Athma A. Pai, Joseph K. Pickrell, Everlyne Nkadori, Yoav Gilad, Jonathan K. Pritchard |
Bioinform. | 2 |
| 2006 | BioHMM: a heterogeneous hidden Markov model for segmenting array CGH dataabstractAbstract Summary: We have developed a new method (BioHMM) for segmenting array comparative genomic hybridization data into states with the same underlying copy number. By utilizing a heterogeneous hidden Markov model, BioHMM incorporates relevant biological factors (e.g. the distance between adjacent clones) in the segmentation process. Availability: BioHMM is available as part of the R library snapCGH which can be downloaded from Contact: [email protected] Supplementary information: Supplementary information is available at John C. Marioni, Natalie P. Thorne, Simon Tavaré |
Bioinform. | 1 |