EDBT 2026 Demo / reviewers in the wild / expert
Marcel H. Schulz
dblp:68/6157
· DBLP profile ↗
29ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-1252-3656ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Predicting gene-specific regulation with transcriptomic and epigenetic single-cell dataabstractMOTIVATION: Analysis of single cell ATAC-seq and RNA-seq data has allowed to gain unprecedented insights into gene regulation by allowing to define cell type-specific regulatory regions and their effects on gene expression. While powerful, such analysis is challenging due to the inherent sparsity of single cell data. RESULTS: We present a new approach, MetaFR, to learn gene-specific models that link open-chromatin variation from scATAC-seq data to gene expression from scRNA-seq. Using efficient regression trees, we illustrate that accurate expression prediction models can be learned on the single-cell or meta-cell level. Validation was done using fine-mapped eQTLs. Meta-cell models were found to outperform single-cell models for most genes. Comparison to the SOTA method SCARlink revealed advantages of MetaFR in terms of runtime and prediction performance. MetaFR thus allows time-efficient analysis and obtains reliable models of gene expression prediction, which can be used to study gene regulation in any organism for which scRNA-seq and scATAC-seq data is available. AVAILABILITY AND IMPLEMENTATION: MetaFR is available under https://github.com/SchulzLab/MetaFR. Laura Rumpf, Fatemeh Behjati-Ardakani, Dennis Hecker, Marcel H. Schulz |
Bioinform. | 4 |
| 2025 | Decoding heart failure subtypes with neural networks via differential explanation analysisabstractSingle-cell transcriptomics offers critical insights into the molecular mechanisms of heart failure (HF) with reduced or preserved ejection fraction. However, understanding these mechanisms is hindered by the growing complexity of single-cell data and the difficulty in unmasking meaningful differential gene signatures among HF types. Machine learning, particularly deep neural networks (NNs), address these challenges by learning transcriptional patterns, reconstructing expression profiles and effectively classifying cells but often lacks interpretability. Recent advances in explainable AI (XAI) offer tools to clarify model decisions. Yet pinpointing differentially regulated genes with these tools remains challenging. We introduce a novel method to identify differentially explained genes (DXGs) based on importance scores derived from custom-built NNs. We highlight the superiority of DXGs in identifying HF subtypes-specific pathways that provide new insights into different types of HF. Offering a robust foundation for future research and therapeutic exploration in expanding transcriptome atlases. Mariano Ruz Jurado, David Rodriguez Morales, Elijah Genetzakis, Fatemeh Behjati-Ardakani, Lukas Zanders, Ariane Fischer, Florian Buettner 0001, Marcel H. Schulz, Stefanie Dimmeler, David John |
Briefings Bioinform. | 8 |
| 2025 | TripLexicon: prediction and analysis of gene regulatory RNA-DNA interactionsabstractMOTIVATION: Non-coding RNA (ncRNA) plays a crucial role in gene regulation, including by forming sequence-specific RNA-DNA interactions at gene regulatory elements. One form of interaction takes place via the formation of RNA:DNA:DNA triple helices (triplexes). Accurate computational prediction of triplex formation from nucleotide sequences is an important tool in ncRNA research but remains somewhat inaccessible and complex. To address this, we created TripLexicon, a web-based interface for accessing and analyzing predicted gene regulatory RNA-DNA interactions in human and mouse. RESULTS: Predicted interactions can be accessed from RNA-, DNA-, and region-centric perspectives. For each RNA transcript, visualizations at genome and nucleotide resolution are available, providing insight into target genes and regions, as well as putative functional domains of the transcript. Predicted target genes can immediately be subjected to ontology and pathway enrichment analysis, providing rapid insight into potential functions mediated by the RNA-DNA interactions of the queried transcript. DNA and region queries are designed to identify potentially important ncRNA interactors at sites of interest. AVAILABILITY AND IMPLEMENTATION: TripLexicon is accessible at https://triplexicon.uni-frankfurt.de. This website is free and open to all users and there is no login requirement. All data and code is uploaded to Zenodo: https://zenodo.org/records/17143608 and the code for the webserver is available on Github: https://github.com/SchulzLab/TripLexicon. Timothy Warwick, Christina Kalk, Ralf P. Brandes, Marcel H. Schulz |
Bioinform. | 4 |
| 2025 | GeneCOCOA: Detecting context-specific functions of individual genes using co-expression dataabstractExtraction of meaningful biological insight from gene expression profiling often focuses on the identification of statistically enriched terms or pathways. These methods typically use gene sets as input data, and subsequently return overrepresented terms along with associated statistics describing their enrichment. This approach does not cater to analyses focused on a single gene-of-interest, particularly when the gene lacks prior functional characterization. To address this, we formulated GeneCOCOA, a method which utilizes context-specific gene co-expression and curated functional gene sets, but focuses on a user-supplied gene-of-interest (GOI). The co-expression between the GOI and subsets of genes from functional groups (e.g. pathways, GO terms) is derived using linear regression, and resulting root-mean-square error values are compared against background values obtained from randomly selected genes. The resulting p values provide a statistical ranking of functional gene sets from any collection, along with their associated terms, based on their co-expression with the gene of interest in a manner specific to the context and experiment. GeneCOCOA thereby provides biological insight into both gene function, and putative regulatory mechanisms by which the expression of the GOI is controlled. Despite its relative simplicity, GeneCOCOA outperforms similar methods in the accurate recall of known gene-disease associations. We furthermore include a differential GeneCOCOA mode, thus presenting the first implementation of a gene-focused approach to experiment-specific gene set enrichment analysis. GeneCOCOA is formulated as an R package for ease-of-use, available at https://github.com/si-ze/geneCOCOA. Simonida Zehr, Thomas Oellerich, Matthias S. Leisegang, Ralf P. Brandes, Marcel H. Schulz, Timothy Warwick |
PLoS Comput. Biol. | 6 |
| 2023 | Debugging Low Power Analog Neural Networks for Edge ComputingabstractIn this paper we present a method to debug and analyze large synthesized ANNs enabling a systematic comparison of the transistor netlist, behavioral model and the implementation. With that an insight into the behavior of the analog netlist is easily gained and errors during generation or badly designed cells are quickly uncovered. An overall judgement of the accuracy is also presented. We demonstrate the functionality on several examples from small ANNs to ANNs consisting of more than 10000 of cells implementing a medical application. Sascha Schmalhofer, Marwin Möller, Nikoletta Katsaouni, Marcel H. Schulz, Lars Hedrich |
DATE | 4 |
| 2023 | Efficiently quantifying DNA methylation for bulk- and single-cell bisulfite dataabstractMOTIVATION: DNA CpG methylation (CpGm) has proven to be a crucial epigenetic factor in the mammalian gene regulatory system. Assessment of DNA CpG methylation values via whole-genome bisulfite sequencing (WGBS) is, however, computationally extremely demanding. RESULTS: We present FAst MEthylation calling (FAME), the first approach to quantify CpGm values directly from bulk or single-cell WGBS reads without intermediate output files. FAME is very fast but as accurate as standard methods, which first produce BS alignment files before computing CpGm values. We present experiments on bulk and single-cell bisulfite datasets in which we show that data analysis can be significantly sped-up and help addressing the current WGBS analysis bottleneck for large-scale datasets without compromising accuracy. AVAILABILITY AND IMPLEMENTATION: An implementation of FAME is open source and licensed under GPL-3.0 at https://github.com/FischerJo/FAME. Jonas Fischer, Marcel H. Schulz |
Bioinform. | 2 |
| 2023 | The adapted Activity-By-Contact model for enhancer-gene assignment and its application to single-cell dataabstractMOTIVATION: Identifying regulatory regions in the genome is of great interest for understanding the epigenomic landscape in cells. One fundamental challenge in this context is to find the target genes whose expression is affected by the regulatory regions. A recent successful method is the Activity-By-Contact (ABC) model which scores enhancer-gene interactions based on enhancer activity and the contact frequency of an enhancer to its target gene. However, it describes regulatory interactions entirely from a gene's perspective, and does not account for all the candidate target genes of an enhancer. In addition, the ABC model requires two types of assays to measure enhancer activity, which limits the applicability. Moreover, there is neither implementation available that could allow for an integration with transcription factor (TF) binding information nor an efficient analysis of single-cell data. RESULTS: We demonstrate that the ABC score can yield a higher accuracy by adapting the enhancer activity according to the number of contacts the enhancer has to its candidate target genes and also by considering all annotated transcription start sites of a gene. Further, we show that the model is comparably accurate with only one assay to measure enhancer activity. We combined our generalized ABC model with TF binding information and illustrated an analysis of a single-cell ATAC-seq dataset of the human heart, where we were able to characterize cell type-specific regulatory interactions and predict gene expression based on TF affinities. All executed processing steps are incorporated into our new computational pipeline STARE. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/schulzlab/STARE. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dennis Hecker, Fatemeh Behjati-Ardakani, Alexander Karollus, Julien Gagneur, Marcel H. Schulz |
Bioinform. | 5 |
| 2022 | A universal model of RNA.DNA: DNA triplex formation accurately predicts genome-wide RNA-DNA interactionsabstractRNA.DNA:DNA triple helix (triplex) formation is a form of RNA-DNA interaction which regulates gene expression but is difficult to study experimentally in vivo. This makes accurate computational prediction of such interactions highly important in the field of RNA research. Current predictive methods use canonical Hoogsteen base pairing rules, which whilst biophysically valid, may not reflect the plastic nature of cell biology. Here, we present the first optimization approach to learn a probabilistic model describing RNA-DNA interactions directly from motifs derived from triplex sequencing data. We find that there are several stable interaction codes, including Hoogsteen base pairing and novel RNA-DNA base pairings, which agree with in vitro measurements. We implemented these findings in TriplexAligner, a program that uses the determined interaction codes to predict triplex binding. TriplexAligner predicts RNA-DNA interactions identified in all-to-all sequencing data more accurately than all previously published tools in human and mouse and also predicts previously studied triplex interactions with known regulatory functions. We further validated a novel triplex interaction using biophysical experiments. Our work is an important step towards better understanding of triplex formation and allows genome-wide analyses of RNA-DNA interactions. Timothy Warwick, Sandra Seredinski, Nina M. Krause, Jasleen Kaur Bains, Lara Althaus, James A Oo, Alessandro Bonetti, Anne Dueck, Stefan Engelhardt, Harald Schwalbe, Matthias S. Leisegang, Marcel H. Schulz, Ralf P. Brandes |
Briefings Bioinform. | 12 |
| 2021 | Fast detection of differential chromatin domains with SCIDDOabstractMOTIVATION: The generation of genome-wide maps of histone modifications using chromatin immunoprecipitation sequencing is a standard approach to dissect the complexity of the epigenome. Interpretation and differential analysis of histone datasets remains challenging due to regulatory meaningful co-occurrences of histone marks and their difference in genomic spread. To ease interpretation, chromatin state segmentation maps are a commonly employed abstraction combining individual histone marks. We developed the tool SCIDDO as a fast, flexible and statistically sound method for the differential analysis of chromatin state segmentation maps. RESULTS: We demonstrate the utility of SCIDDO in a comparative analysis that identifies differential chromatin domains (DCD) in various regulatory contexts and with only moderate computational resources. We show that the identified DCDs correlate well with observed changes in gene expression and can recover a substantial number of differentially expressed genes (DEGs). We showcase SCIDDO's ability to directly interrogate chromatin dynamics, such as enhancer switches in downstream analysis, which simplifies exploring specific questions about regulatory changes in chromatin. By comparing SCIDDO to competing methods, we provide evidence that SCIDDO's performance in identifying DEGs via differential chromatin marking is more stable across a range of cell-type comparisons and parameter cut-offs. AVAILABILITY AND IMPLEMENTATION: The SCIDDO source code is openly available under github.com/ptrebert/sciddo. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peter Ebert, Marcel H. Schulz |
Bioinform. | 2 |
| 2021 | Fostering accessible online education using Galaxy as an e-learning platformabstractThe COVID-19 pandemic is shifting teaching to an online setting all over the world. The Galaxy framework facilitates the online learning process and makes it accessible by providing a library of high-quality community-curated training materials, enabling easy access to data and tools, and facilitates sharing achievements and progress between students and instructors. By combining Galaxy with robust communication channels, effective instruction can be designed inclusively, regardless of the students' environments. Beatriz Serrano-Solano, Melanie Christine Föll, Cristóbal Gallardo-Alba, Anika Erxleben-Eggenhofer, Helena Rasche, Saskia D. Hiltemann, Matthias Fahrner, Mark J. Dunning, Marcel H. Schulz, Beáta Scholtz, Dave Clements, Anton Nekrutenko, Bérénice Batut, Björn A. Grüning |
PLoS Comput. Biol. | 9 |
| 2020 | Improved linking of motifs to their TFs using domain informationabstractMOTIVATION: A central aim of molecular biology is to identify mechanisms of transcriptional regulation. Transcription factors (TFs), which are DNA-binding proteins, are highly involved in these processes, thus a crucial information is to know where TFs interact with DNA and to be aware of the TFs' DNA-binding motifs. For that reason, computational tools exist that link DNA-binding motifs to TFs either without sequence information or based on TF-associated sequences, e.g. identified via a chromatin immunoprecipitation followed by sequencing (ChIP-seq) experiment.In this paper, we present MASSIF, a novel method to improve the performance of existing tools that link motifs to TFs relying on TF-associated sequences. MASSIF is based on the idea that a DNA-binding motif, which is correctly linked to a TF, should be assigned to a DNA-binding domain (DBD) similar to that of the mapped TF. Because DNA-binding motifs are in general not linked to DBDs, it is not possible to compare the DBD of a TF and the motif directly. Instead we created a DBD collection, which consist of TFs with a known DBD and an associated motif. This collection enables us to evaluate how likely it is that a linked motif and a TF of interest are associated to the same DBD. We named this similarity measure domain score, and represent it as a P-value. We developed two different ways to improve the performance of existing tools that link motifs to TFs based on TF-associated sequences: (i) using meta-analysis to combine P-values from one or several of these tools with the P-value of the domain score and (ii) filter unlikely motifs based on the domain score. RESULTS: We demonstrate the functionality of MASSIF on several human ChIP-seq datasets, using either motifs from the HOCOMOCO database or de novo identified ones as input motifs. In addition, we show that both variants of our method improve the performance of tools that link motifs to TFs based on TF-associated sequences significantly independent of the considered DBD type. AVAILABILITY AND IMPLEMENTATION: MASSIF is freely available online at https://github.com/SchulzLab/MASSIF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nina Baumgarten, Florian Schmidt 0003, Marcel H. Schulz |
Bioinform. | 3 |
| 2019 | Large-scale inference of competing endogenous RNA networks with sparse partial correlationabstractMOTIVATION: MicroRNAs (miRNAs) are important non-coding post-transcriptional regulators that are involved in many biological processes and human diseases. Individual miRNAs may regulate hundreds of genes, giving rise to a complex gene regulatory network in which transcripts carrying miRNA binding sites act as competing endogenous RNAs (ceRNAs). Several methods for the analysis of ceRNA interactions exist, but these do often not adjust for statistical confounders or address the problem that more than one miRNA interacts with a target transcript. RESULTS: We present SPONGE, a method for the fast construction of ceRNA networks. SPONGE uses 'multiple sensitivity correlation', a newly defined measure for which we can estimate a distribution under a null hypothesis. SPONGE can accurately quantify the contribution of multiple miRNAs to a ceRNA interaction with a probabilistic model that addresses previously neglected confounding factors and allows fast P-value calculation, thus outperforming existing approaches. We applied SPONGE to paired miRNA and gene expression data from The Cancer Genome Atlas for studying global effects of miRNA-mediated cross-talk. Our results highlight already established and novel protein-coding and non-coding ceRNAs which could serve as biomarkers in cancer. AVAILABILITY AND IMPLEMENTATION: SPONGE is available as an R/Bioconductor package (doi: 10.18129/B9.bioc.SPONGE). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Markus List, Azim Dehghani Amirabad, Dennis Kostka, Marcel H. Schulz |
Bioinform. | 4 |
| 2019 | TEPIC 2 - an extended framework for transcription factor binding prediction and integrative epigenomic analysisabstractSUMMARY: Prediction of transcription factor (TF) binding from epigenetics data and integrative analysis thereof are challenging. Here, we present TEPIC 2 a framework allowing for fast, accurate and versatile prediction, and analysis of TF binding from epigenetics data: it supports 30 species with binding motifs, computes TF gene and scores up to two orders of magnitude faster than before due to improved implementation, and offers easy-to-use machine learning pipelines for integrated analysis of TF binding predictions with gene expression data allowing the identification of important TFs. AVAILABILITY AND IMPLEMENTATION: TEPIC is implemented in C++, R, and Python. It is freely available at https://github.com/SchulzLab/TEPIC and can be used on Linux based systems. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Florian Schmidt 0003, Fabian Kern, Peter Ebert, Nina Baumgarten, Marcel H. Schulz |
Bioinform. | 5 |
| 2019 | On the problem of confounders in modeling gene expressionabstractMOTIVATION: Modeling of Transcription Factor (TF) binding from both ChIP-seq and chromatin accessibility data has become prevalent in computational biology. Several models have been proposed to generate new hypotheses on transcriptional regulation. However, there is no distinct approach to derive TF binding scores from ChIP-seq and open chromatin experiments. Here, we review biases of various scoring approaches and their effects on the interpretation and reliability of predictive gene expression models. RESULTS: We generated predictive models for gene expression using ChIP-seq and DNase1-seq data from DEEP and ENCODE. Via randomization experiments, we identified confounders in TF gene scores derived from both ChIP-seq and DNase1-seq data. We reviewed correction approaches for both data types, which reduced the influence of identified confounders without harm to model performance. Also, our analyses highlighted further quality control measures, in addition to model performance, that may help to assure model reliability and to avoid misinterpretation in future studies. AVAILABILITY AND IMPLEMENTATION: The software used in this study is available online at https://github.com/SchulzLab/TEPIC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Florian Schmidt 0003, Marcel H. Schulz |
Bioinform. | 2 |
| 2018 | Chromatyping: Reconstructing Nucleosome Profiles from NOMe Sequencing Data
Shounak Chakraborty 0003, Stefan Canzar, Tobias Marschall, Marcel H. Schulz |
RECOMB | 4 |
| 2018 | An ontology-based method for assessing batch effect adjustment approaches in heterogeneous datasetsabstractMotivation: International consortia such as the Genotype-Tissue Expression (GTEx) project, The Cancer Genome Atlas (TCGA) or the International Human Epigenetics Consortium (IHEC) have produced a wealth of genomic datasets with the goal of advancing our understanding of cell differentiation and disease mechanisms. However, utilizing all of these data effectively through integrative analysis is hampered by batch effects, large cell type heterogeneity and low replicate numbers. To study if batch effects across datasets can be observed and adjusted for, we analyze RNA-seq data of 215 samples from ENCODE, Roadmap, BLUEPRINT and DEEP as well as 1336 samples from GTEx and TCGA. While batch effects are a considerable issue, it is non-trivial to determine if batch adjustment leads to an improvement in data quality, especially in cases of low replicate numbers. Results: We present a novel method for assessing the performance of batch effect adjustment methods on heterogeneous data. Our method borrows information from the Cell Ontology to establish if batch adjustment leads to a better agreement between observed pairwise similarity and similarity of cell types inferred from the ontology. A comparison of state-of-the art batch effect adjustment methods suggests that batch effects in heterogeneous datasets with low replicate numbers cannot be adequately adjusted. Better methods need to be developed, which can be assessed objectively in the framework presented here. Availability and implementation: Our method is available online at https://github.com/SchulzLab/OntologyEval. Supplementary information: Supplementary data are available at Bioinformatics online. Florian Schmidt 0003, Markus List, Engin Cukuroglu, Sebastian Köhler 0001, Jonathan Göke, Marcel H. Schulz |
Bioinform. | 6 |
| 2018 | In silico read normalization using set multi-cover optimizationabstractMotivation: De Bruijn graphs are a common assembly data structure for sequencing datasets. But with the advances in sequencing technologies, assembling high coverage datasets has become a computational challenge. Read normalization, which removes redundancy in datasets, is widely applied to reduce resource requirements. Current normalization algorithms, though efficient, provide no guarantee to preserve important k-mers that form connections between regions in the graph. Results: Here, normalization is phrased as a set multi-cover problem on reads and a heuristic algorithm, Optimized Read Normalization Algorithm (ORNA), is proposed. ORNA normalizes to the minimum number of reads required to retain all k-mers and their relative k-mer abundances from the original dataset. Hence, all connections from the original graph are preserved. ORNA was tested on various RNA-seq datasets with different coverage values. It was compared to the current normalization algorithms and was found to be performing better. Normalizing error corrected data allows for more accurate assemblies compared to the normalized uncorrected dataset. Further, an application is proposed in which multiple datasets are combined and normalized to predict novel transcripts that would have been missed otherwise. Finally, ORNA is a general purpose normalization algorithm that is fast and significantly reduces datasets with loss of assembly quality in between [1, 30]% depending on reduction stringency. Availability and implementation: ORNA is available at https://github.com/SchulzLab/ORNA. Supplementary information: Supplementary data are available at Bioinformatics online. Dilip A. Durai, Marcel H. Schulz |
Bioinform. | 2 |
| 2018 | JAMI: fast computation of conditional mutual information for ceRNA network analysisabstractMotivation: Genome-wide measurements of paired miRNA and gene expression data have enabled the prediction of competing endogenous RNAs (ceRNAs). It has been shown that the sponge effect mediated by protein-coding as well as non-coding ceRNAs can play an important regulatory role in the cell in health and disease. Therefore, many computational methods for the computational identification of ceRNAs have been suggested. In particular, methods based on Conditional Mutual Information (CMI) have shown promising results. However, the currently available implementation is slow and cannot be used to perform computations on a large scale. Results: Here, we present JAMI, a Java tool that uses a non-parametric estimator for CMI values from gene and miRNA expression data. We show that JAMI speeds up the computation of ceRNA networks by a factor of ∼70 compared to currently available implementations. Further, JAMI supports multi-threading to make use of common multi-core architectures for further performance gain. Requirements: Java 8. Availability and implementation: JAMI is available as open-source software from https://github.com/SchulzLab/JAMI. Supplementary information: Supplementary data are available at Bioinformatics online. Andrea Hornáková, Markus List, Jilles Vreeken, Marcel H. Schulz |
Bioinform. | 4 |
| 2016 | Informed kmer selection for de novo transcriptome assemblyabstractMOTIVATION: De novo transcriptome assembly is an integral part for many RNA-seq workflows. Common applications include sequencing of non-model organisms, cancer or meta transcriptomes. Most de novo transcriptome assemblers use the de Bruijn graph (DBG) as the underlying data structure. The quality of the assemblies produced by such assemblers is highly influenced by the exact word length k As such no single kmer value leads to optimal results. Instead, DBGs over different kmer values are built and the assemblies are merged to improve sensitivity. However, no studies have investigated thoroughly the problem of automatically learning at which kmer value to stop the assembly. Instead a suboptimal selection of kmer values is often used in practice. RESULTS: Here we investigate the contribution of a single kmer value in a multi-kmer based assembly approach. We find that a comparative clustering of related assemblies can be used to estimate the importance of an additional kmer assembly. Using a model fit based algorithm we predict the kmer value at which no further assemblies are necessary. Our approach is tested with different de novo assemblers for datasets with different coverage values and read lengths. Further, we suggest a simple post processing step that significantly improves the quality of multi-kmer assemblies. CONCLUSION: We provide an automatic method for limiting the number of kmer values without a significant loss in assembly quality but with savings in assembly time. This is a step forward to making multi-kmer methods more reliable and easier to use. AVAILABILITY AND IMPLEMENTATION: A general implementation of our approach can be found under: https://github.com/SchulzLab/KREATIONSupplementary information: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Dilip A. Durai, Marcel H. Schulz |
Bioinform. | 2 |
| 2014 | Fiona: a parallel and automatic strategy for read error correctionabstractMOTIVATION: Automatic error correction of high-throughput sequencing data can have a dramatic impact on the amount of usable base pairs and their quality. It has been shown that the performance of tasks such as de novo genome assembly and SNP calling can be dramatically improved after read error correction. While a large number of methods specialized for correcting substitution errors as found in Illumina data exist, few methods for the correction of indel errors, common to technologies like 454 or Ion Torrent, have been proposed. RESULTS: We present Fiona, a new stand-alone read error-correction method. Fiona provides a new statistical approach for sequencing error detection and optimal error correction and estimates its parameters automatically. Fiona is able to correct substitution, insertion and deletion errors and can be applied to any sequencing technology. It uses an efficient implementation of the partial suffix array to detect read overlaps with different seed lengths in parallel. We tested Fiona on several real datasets from a variety of organisms with different read lengths and compared its performance with state-of-the-art methods. Fiona shows a constantly higher correction accuracy over a broad range of datasets from 454 and Ion Torrent sequencers, without compromise in speed. CONCLUSION: Fiona is an accurate parameter-free read error-correction method that can be run on inexpensive hardware and can make use of multicore parallelization whenever available. Fiona was implemented using the SeqAn library for sequence analysis and is publicly available for download at http://www.seqan.de/projects/fiona. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marcel H. Schulz, David Weese, Manuel Holtgrewe, Viktoria Dimitrova, Sijia Niu, Knut Reinert, Hugues Richard |
Bioinform. | 1 |
| 2012 | Bayesian ontology querying for accurate and noise-tolerant semantic searchesabstractMOTIVATION: Ontologies provide a structured representation of the concepts of a domain of knowledge as well as the relations between them. Attribute ontologies are used to describe the characteristics of the items of a domain, such as the functions of proteins or the signs and symptoms of disease, which opens the possibility of searching a database of items for the best match to a list of observed or desired attributes. However, naive search methods do not perform well on realistic data because of noise in the data, imprecision in typical queries and because individual items may not display all attributes of the category they belong to. RESULTS: We present a method for combining ontological analysis with Bayesian networks to deal with noise, imprecision and attribute frequencies and demonstrate an application of our method as a differential diagnostic support system for human genetics. AVAILABILITY: We provide an implementation for the algorithm and the benchmark at http://compbio.charite.de/boqa/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary Material for this article is available at Bioinformatics online. Sebastian Bauer 0002, Sebastian Köhler 0001, Marcel H. Schulz, Peter N. Robinson |
Bioinform. | 3 |
| 2012 | Detecting genomic indel variants with exact breakpoints in single- and paired-end sequencing data using SplazerSabstractMOTIVATION: The reliable detection of genomic variation in resequencing data is still a major challenge, especially for variants larger than a few base pairs. Sequencing reads crossing boundaries of structural variation carry the potential for their identification, but are difficult to map. RESULTS: Here we present a method for 'split' read mapping, where prefix and suffix match of a read may be interrupted by a longer gap in the read-to-reference alignment. We use this method to accurately detect medium-sized insertions and long deletions with precise breakpoints in genomic resequencing data. Compared with alternative split mapping methods, SplazerS significantly improves sensitivity for detecting large indel events, especially in variant-rich regions. Our method is robust in the presence of sequencing errors as well as alignment errors due to genomic mutations/divergence, and can be used on reads of variable lengths. Our analysis shows that SplazerS is a versatile tool applicable to unanchored or single-end as well as anchored paired-end reads. In addition, application of SplazerS to targeted resequencing data led to the interesting discovery of a complete, possibly functional gene retrocopy variant. AVAILABILITY: SplazerS is available from http://www.seqan.de/projects/ splazers. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anne-Katrin Emde, Marcel H. Schulz, David Weese, Ruping Sun, Martin Vingron, Vera M. Kalscheuer, Stefan A. Haas, Knut Reinert |
Bioinform. | 2 |
| 2012 | Estimation of pairwise sequence similarity of mammalian enhancers with word neighbourhood countsabstractMOTIVATION: The identity of cells and tissues is to a large degree governed by transcriptional regulation. A major part is accomplished by the combinatorial binding of transcription factors at regulatory sequences, such as enhancers. Even though binding of transcription factors is sequence-specific, estimating the sequence similarity of two functionally similar enhancers is very difficult. However, a similarity measure for regulatory sequences is crucial to detect and understand functional similarities between two enhancers and will facilitate large-scale analyses like clustering, prediction and classification of genome-wide datasets. RESULTS: We present the standardized alignment-free sequence similarity measure N2, a flexible framework that is defined for word neighbourhoods. We explore the usefulness of adding reverse complement words as well as words including mismatches into the neighbourhood. On simulated enhancer sequences as well as functional enhancers in mouse development, N2 is shown to outperform previous alignment-free measures. N2 is flexible, faster than competing methods and less susceptible to single sequence noise and the occurrence of repetitive sequences. Experiments on the mouse enhancers reveal that enhancers active in different tissues can be separated by pairwise comparison using N2. CONCLUSION: N2 represents an improvement over previous alignment-free similarity measures without compromising speed, which makes it a good candidate for large-scale sequence comparison of regulatory sequences. AVAILABILITY: The software is part of the open-source C++ library SeqAn (www.seqan.de) and a compiled version can be downloaded at http://www.seqan.de/projects/alf.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonathan Göke, Marcel H. Schulz, Julia Lasserre, Martin Vingron |
Bioinform. | 2 |
| 2012 | Oases: robust de novo RNA-seq assembly across the dynamic range of expression levelsabstractMOTIVATION: High-throughput sequencing has made the analysis of new model organisms more affordable. Although assembling a new genome can still be costly and difficult, it is possible to use RNA-seq to sequence mRNA. In the absence of a known genome, it is necessary to assemble these sequences de novo, taking into account possible alternative isoforms and the dynamic range of expression values. RESULTS: We present a software package named Oases designed to heuristically assemble RNA-seq reads in the absence of a reference genome, across a broad spectrum of expression values and in presence of alternative isoforms. It achieves this by using an array of hash lengths, a dynamic filtering of noise, a robust resolution of alternative splicing events and the efficient merging of multiple assemblies. It was tested on human and mouse RNA-seq data and is shown to improve significantly on the transABySS and Trinity de novo transcriptome assemblers. AVAILABILITY AND IMPLEMENTATION: Oases is freely available under the GPL license at www.ebi.ac.uk/~zerbino/oases/. Marcel H. Schulz, Daniel R. Zerbino, Martin Vingron, Ewan Birney |
Bioinform. | 1 |
| 2011 | DECOD: fast and accurate discriminative DNA motif findingabstractMOTIVATION: Motif discovery is now routinely used in high-throughput studies including large-scale sequencing and proteomics. These datasets present new challenges. The first is speed. Many motif discovery methods do not scale well to large datasets. Another issue is identifying discriminative rather than generative motifs. Such discriminative motifs are important for identifying co-factors and for explaining changes in behavior between different conditions. RESULTS: To address these issues we developed a method for DECOnvolved Discriminative motif discovery (DECOD). DECOD uses a k-mer count table and so its running time is independent of the size of the input set. By deconvolving the k-mers DECOD considers context information without using the sequences directly. DECOD outperforms previous methods both in speed and in accuracy when using simulated and real biological benchmark data. We performed new binding experiments for p53 mutants and used DECOD to identify p53 co-factors, suggesting new mechanisms for p53 activation. AVAILABILITY: The source code and binaries for DECOD are available at http://www.sb.cs.cmu.edu/DECOD CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peter Huggins, Idit Shiff, Rachel Beckerman, Oleg Laptenko, Carol Prives, Marcel H. Schulz, Itamar Simon, Ziv Bar-Joseph |
Bioinform. | 7 |
| 2011 | Exact score distribution computation for ontological similarity searchesabstractBACKGROUND: Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., finding functionally related proteins with the Gene Ontology or phenotypically similar diseases with the Human Phenotype Ontology (HPO). We have recently shown that the performance of semantic similarity searches can be improved by ranking results according to the probability of obtaining a given score at random rather than by the scores themselves. However, to date, there are no algorithms for computing the exact distribution of semantic similarity scores, which is necessary for computing the exact P-value of a given score. RESULTS: In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a simple null hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the HPO. It is shown that exact P-value calculation improves clinical diagnosis using the HPO compared to approaches based on sampling. CONCLUSIONS: The new algorithm enables for the first time exact P-value calculation via exact score distribution computation for ontology similarity searches. The approach is applicable to any ontology for which the annotation-propagation rule holds and can improve any bioinformatic method that makes only use of the raw similarity scores. The algorithm was implemented in Java, supports any ontology in OBO format, and is available for non-commercial and academic usage under: https://compbio.charite.de/svn/hpo/trunk/src/tools/significance/ Marcel H. Schulz, Sebastian Köhler 0001, Sebastian Bauer 0002, Peter N. Robinson |
BMC Bioinform. | 1 |
| 2009 | Exact Score Distribution Computation for Similarity Searches in Ontologies
Marcel H. Schulz, Sebastian Köhler 0001, Sebastian Bauer 0002, Martin Vingron, Peter N. Robinson |
WABI | 1 |
| 2009 | Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short readsabstractMOTIVATION: There is a strong demand in the genomic community to develop effective algorithms to reliably identify genomic variants. Indel detection using next-gen data is difficult and identification of long structural variations is extremely challenging. RESULTS: We present Pindel, a pattern growth approach, to detect breakpoints of large deletions and medium-sized insertions from paired-end short reads. We use both simulated reads and real data to demonstrate the efficiency of the computer program and accuracy of the results. AVAILABILITY: The binary code and a short user manual can be freely downloaded from http://www.ebi.ac.uk/ approximately kye/pindel/. CONTACT: [email protected]; [email protected]. Kai Ye 0001, Marcel H. Schulz, Rolf Apweiler, Zemin Ning |
Bioinform. | 2 |
| 2008 | Fast and Adaptive Variable Order Markov Chain Construction
Marcel H. Schulz, David Weese, Tobias Rausch, Andreas Gogol-Döring, Knut Reinert, Martin Vingron |
WABI | 1 |