Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jonathan Göke

dblp:88/10995 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-0825-4991ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis
nanopore sequencing
0.912025
Leveraging basecaller's move table to generate a lightweight k-mer model for nanopore sequencing analysis · Bioinform. 2025
Bioinformatics and computational biology › sequence analysis › pattern matching
approximate string matching
0.812024
Flexiplex: a versatile demultiplexer and search tool for omics data · Bioinform. 2024
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.812024
Flexiplex: a versatile demultiplexer and search tool for omics data · Bioinform. 2024
Bioinformatics and computational biology › omics data analysis
batch effect correction
0.312018
An ontology-based method for assessing batch effect adjustment approaches in heterogeneous datasets · Bioinform. 2018
Bioinformatics and computational biology › single-cell analysis › single-cell RNA sequencing
single-cell RNA-seq analysis
0.212024
Flexiplex: a versatile demultiplexer and search tool for omics data · Bioinform. 2024
Bioinformatics and computational biology › sequence analysis › sequence comparison
alignment-free sequence comparison
0.112012
Estimation of pairwise sequence similarity of mammalian enhancers with word neighbourhood counts · Bioinform. 2012
Bioinformatics and computational biology › gene regulation
regulatory genomics
0.112012
Estimation of pairwise sequence similarity of mammalian enhancers with word neighbourhood counts · Bioinform. 2012
Bioinformatics and computational biology
sequence analysis
0.112012
Estimation of pairwise sequence similarity of mammalian enhancers with word neighbourhood counts · Bioinform. 2012

Methods — techniques the papers use, named apart from their topics

move table analysis · 0.9levenshtein distance · 0.8pairwise similarity · 0.3cell ontology · 0.3word neighbourhood counting · 0.1reverse complement · 0.1
YearPublicationVenuePosition
2025 Leveraging basecaller's move table to generate a lightweight k-mer model for nanopore sequencing analysis
abstract
MOTIVATION: Nanopore sequencing by Oxford Nanopore Technologies (ONT) enables direct analysis of DNA and RNA by capturing raw electrical signals. Different nanopore chemistries have varied k-mer lengths, current levels, and standard deviations, which are stored in "k-mer models." In cases where official models are lacking or unsuitable for specific sequencing conditions, tailored k-mer models are crucial to ensure precise signal-to-sequence alignment, analysis and interpretation. The process of transforming raw signal data into nucleotide sequences, known as basecalling, is a fundamental step in nanopore sequencing. RESULTS: In this study, we leverage the move table produced by ONT's basecalling software to create a lightweight de novo k-mer model for RNA004 chemistry. We demonstrate the validity of our custom k-mer model by using it to guide signal-to-sequence alignment analysis, achieving high alignment rates (97.48%) compared to larger default models. Additionally, our 5-mer model exhibits similar performance as the default 9-mer models another analysis, such as detection of m6A RNA modifications. We provide our method, termed Poregen, as a generalizable approach for creation of custom, de novo k-mer models for nanopore signal data analysis. AVAILABILITY AND IMPLEMENTATION: Poregen is an open source package under an MIT license: https://github.com/hiruna72/poregen.
Hiruna Samarakoon, Yuk Kei Wan, Sri Parameswaran, Jonathan Göke, Hasindu Gamaarachchi, Ira W. Deveson
Bioinform.4
2024 Flexiplex: a versatile demultiplexer and search tool for omics data
abstract
MOTIVATION: The process of analyzing high throughput sequencing data often requires the identification and extraction of specific target sequences. This could include tasks, such as identifying cellular barcodes and UMIs in single-cell data, and specific genetic variants for genotyping. However, existing tools, which perform these functions are often task-specific, such as only demultiplexing barcodes for a dedicated type of experiment, or are not tolerant to noise in the sequencing data. RESULTS: To overcome these limitations, we developed Flexiplex, a versatile and fast sequence searching and demultiplexing tool for omics data, which is based on the Levenshtein distance and thus allows imperfect matches. We demonstrate Flexiplex's application on three use cases, identifying cell-line-specific sequences in Illumina short-read single-cell data, and discovering and demultiplexing cellular barcodes from noisy long-read single-cell RNA-seq data. We show that Flexiplex achieves an excellent balance of accuracy and computational efficiency compared to leading task-specific tools. AVAILABILITY AND IMPLEMENTATION: Flexiplex is available at https://davidsongroup.github.io/flexiplex/.
Oliver Cheng, Min Hao Ling, Shuyi Wu, Matthew E. Ritchie, Jonathan Göke, Noorul Amin, Nadia M. Davidson
Bioinform.6
2018 An ontology-based method for assessing batch effect adjustment approaches in heterogeneous datasets
abstract
Motivation: International consortia such as the Genotype-Tissue Expression (GTEx) project, The Cancer Genome Atlas (TCGA) or the International Human Epigenetics Consortium (IHEC) have produced a wealth of genomic datasets with the goal of advancing our understanding of cell differentiation and disease mechanisms. However, utilizing all of these data effectively through integrative analysis is hampered by batch effects, large cell type heterogeneity and low replicate numbers. To study if batch effects across datasets can be observed and adjusted for, we analyze RNA-seq data of 215 samples from ENCODE, Roadmap, BLUEPRINT and DEEP as well as 1336 samples from GTEx and TCGA. While batch effects are a considerable issue, it is non-trivial to determine if batch adjustment leads to an improvement in data quality, especially in cases of low replicate numbers. Results: We present a novel method for assessing the performance of batch effect adjustment methods on heterogeneous data. Our method borrows information from the Cell Ontology to establish if batch adjustment leads to a better agreement between observed pairwise similarity and similarity of cell types inferred from the ontology. A comparison of state-of-the art batch effect adjustment methods suggests that batch effects in heterogeneous datasets with low replicate numbers cannot be adequately adjusted. Better methods need to be developed, which can be assessed objectively in the framework presented here. Availability and implementation: Our method is available online at https://github.com/SchulzLab/OntologyEval. Supplementary information: Supplementary data are available at Bioinformatics online.
Florian Schmidt 0003, Markus List, Engin Cukuroglu, Sebastian Köhler 0001, Jonathan Göke, Marcel H. Schulz
Bioinform.5
2012 Estimation of pairwise sequence similarity of mammalian enhancers with word neighbourhood counts
abstract
MOTIVATION: The identity of cells and tissues is to a large degree governed by transcriptional regulation. A major part is accomplished by the combinatorial binding of transcription factors at regulatory sequences, such as enhancers. Even though binding of transcription factors is sequence-specific, estimating the sequence similarity of two functionally similar enhancers is very difficult. However, a similarity measure for regulatory sequences is crucial to detect and understand functional similarities between two enhancers and will facilitate large-scale analyses like clustering, prediction and classification of genome-wide datasets. RESULTS: We present the standardized alignment-free sequence similarity measure N2, a flexible framework that is defined for word neighbourhoods. We explore the usefulness of adding reverse complement words as well as words including mismatches into the neighbourhood. On simulated enhancer sequences as well as functional enhancers in mouse development, N2 is shown to outperform previous alignment-free measures. N2 is flexible, faster than competing methods and less susceptible to single sequence noise and the occurrence of repetitive sequences. Experiments on the mouse enhancers reveal that enhancers active in different tissues can be separated by pairwise comparison using N2. CONCLUSION: N2 represents an improvement over previous alignment-free similarity measures without compromising speed, which makes it a good candidate for large-scale sequence comparison of regulatory sequences. AVAILABILITY: The software is part of the open-source C++ library SeqAn (www.seqan.de) and a compiled version can be downloaded at http://www.seqan.de/projects/alf.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jonathan Göke, Marcel H. Schulz, Julia Lasserre, Martin Vingron
Bioinform.1
2011 Combinatorial Binding in Human and Mouse Embryonic Stem Cells Identifies Conserved Enhancers Active in Early Embryonic Development
abstract
Transcription factors are proteins that regulate gene expression by binding to cis-regulatory sequences such as promoters and enhancers. In embryonic stem (ES) cells, binding of the transcription factors OCT4, SOX2 and NANOG is essential to maintain the capacity of the cells to differentiate into any cell type of the developing embryo. It is known that transcription factors interact to regulate gene expression. In this study we show that combinatorial binding is strongly associated with co-localization of the transcriptional co-activator Mediator, H3K27ac and increased expression of nearby genes in embryonic stem cells. We observe that the same loci bound by Oct4, Nanog and Sox2 in ES cells frequently drive expression in early embryonic development. Comparison of mouse and human ES cells shows that less than 5% of individual binding events for OCT4, SOX2 and NANOG are shared between species. In contrast, about 15% of combinatorial binding events and even between 53% and 63% of combinatorial binding events at enhancers active in early development are conserved. Our analysis suggests that the combination of OCT4, SOX2 and NANOG binding is critical for transcription in ES cells and likely plays an important role for embryogenesis by binding at conserved early developmental enhancers. Our data suggests that the fast evolutionary rewiring of regulatory networks mainly affects individual binding events, whereas "gene regulatory hotspots" which are bound by multiple factors and active in multiple tissues throughout early development are under stronger evolutionary constraints.
Jonathan Göke, Marc Jung, Sarah Behrens, Lukas Chavez, Sean O'Keeffe, Bernd Timmermann, Hans Lehrach, James Adjaye, Martin Vingron
PLoS Comput. Biol.1