VLDB 2026 Research / reviewers in the wild / expert
Christoph Dieterich
dblp:69/5744
· DBLP profile ↗
16ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0001-9468-6311ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ψ-co-mAFiA: concurrent detection of pseudouridine and m6A in single RNA moleculesabstractSUMMARY: The development of third-generation sequencing technologies enables the detection of RNA modifications at single-molecule resolution. Specifically for direct RNA sequencing on the ONT platform, we have previously developed an m6A detection algorithm called mAFiA. Here, we present the updated method, now covering all 18 DRACH m6A contexts as well as the identification of pseudouridine sites (Ψ). Our modification level predictions compare favorably with orthogonal methods and respond to knockdown or knock out of writer proteins. The simultaneous detection of multiple modifications on a single RNA molecule opens up the possibility to study cross-modification interactions. AVAILABILITY AND IMPLEMENTATION: Ψ-co-mAFiA is available at https://github.com/dieterich-lab/psi-co-mAFiA and licensed under GPLv3.0. An archived version of the software is available on Zenodo at https://doi.org/10.5281/zenodo.16797676. Adrian Chan, Isabel S. Naarmann-de Vries, Christoph Dieterich |
Bioinform. | 3 |
| 2025 | Medication information extraction using local large language modelsabstractOBJECTIVE: Medication information is crucial for clinical routine and research. However, a vast amount is stored in unstructured text, such as doctor's letters, requiring manual extraction - a resource-intensive, error-prone task. Automating this process comes with significant constraints in a clinical setup, including the demand for clinical expertise, limited time-resources, restricted IT infrastructure, and the demand for transparent predictions. Recent advances in generative large language models (LLMs) and parameter-efficient fine-tuning methods show potential to address these challenges. METHODS: We evaluated local LLMs for end-to-end extraction of medication information, combining named entity recognition and relation extraction. We used format-restricting instructions and developed an innovative feedback pipeline to facilitate automated evaluation. We applied token-level Shapley values to visualize and quantify token contributions, to improve transparency of model predictions. RESULTS: Two open-source LLMs - one general (Llama) and one domain-specific (OpenBioLLM) - were evaluated on the English n2c2 2018 corpus and the German CARDIO:DE corpus. OpenBioLLM frequently struggled with structured outputs and hallucinations. Fine-tuned Llama models achieved new state-of-the-art results, improving F1-score by up to 10 percentage points (pp.) for adverse drug events and 6 pp. for medication reasons on English data. On the German dataset, Llama established a new benchmark, outperforming traditional machine learning methods by up to 16 pp. micro average F1-score. CONCLUSION: Our findings show that fine-tuned local open-source generative LLMs outperform SOTA methods for medication information extraction, delivering high performance with limited time and IT resources in a real-world clinical setup, and demonstrate their effectiveness on both English and German data. Applying Shapley values improved prediction transparency, supporting informed clinical decision-making. Phillip Richter-Pechanski, Marvin Seiferling, Christina Kiriakou, Dominic M. Schwab, Nicolas Geis, Christoph Dieterich, Anette Frank |
J. Biomed. Informatics | 6 |
| 2023 | Characterizing alternative splicing effects on protein interaction networks with LINDAabstractMOTIVATION: Alternative RNA splicing plays a crucial role in defining protein function. However, despite its relevance, there is a lack of tools that characterize effects of splicing on protein interaction networks in a mechanistic manner (i.e. presence or absence of protein-protein interactions due to RNA splicing). To fill this gap, we present Linear Integer programming for Network reconstruction using transcriptomics and Differential splicing data Analysis (LINDA) as a method that integrates resources of protein-protein and domain-domain interactions, transcription factor targets, and differential splicing/transcript analysis to infer splicing-dependent effects on cellular pathways and regulatory networks. RESULTS: We have applied LINDA to a panel of 54 shRNA depletion experiments in HepG2 and K562 cells from the ENCORE initiative. Through computational benchmarking, we could show that the integration of splicing effects with LINDA can identify pathway mechanisms contributing to known bioprocesses better than other state of the art methods, which do not account for splicing. Additionally, we have experimentally validated some of the predicted splicing effects that the depletion of HNRNPK in K562 cells has on signalling. Enio Gjerga, Isabel S. Naarmann-de Vries, Christoph Dieterich |
Bioinform. | 3 |
| 2021 | A comparison of metabolic labeling and statistical methods to infer genome-wide dynamics of RNA turnoverabstractMetabolic labeling of newly transcribed RNAs coupled with RNA-seq is being increasingly used for genome-wide analysis of RNA dynamics. Methods including standard biochemical enrichment and recent nucleotide conversion protocols each require special experimental and computational treatment. Despite their immediate relevance, these technologies have not yet been assessed and benchmarked, and no data are currently available to advance reproducible research and the development of better inference tools. Here, we present a systematic evaluation and comparison of four RNA labeling protocols: 4sU-tagging biochemical enrichment, including spike-in RNA controls, SLAM-seq, TimeLapse-seq and TUC-seq. All protocols are evaluated based on practical considerations, conversion efficiency and wet lab requirements to handle hazardous substances. We also compute decay rate estimates and confidence intervals for each protocol using two alternative statistical frameworks, pulseR and GRAND-SLAM, for over 11 600 human genes and evaluate the underlying computational workflows for their robustness and ease of use. Overall, we demonstrate a high inter-method reliability across eight use case scenarios. Our results and data will facilitate reproducible research and serve as a resource contributing to a fuller understanding of RNA biology. Etienne Boileau, Janine Altmüller, Isabel S. Naarmann-de Vries, Christoph Dieterich |
Briefings Bioinform. | 4 |
| 2019 | circtools - a one-stop software solution for circular RNA researchabstractMOTIVATION: Circular RNAs (circRNAs) originate through back-splicing events from linear primary transcripts, are resistant to exonucleases, are not polyadenylated and have been shown to be highly specific for cell type and developmental stage. CircRNA detection starts from high-throughput sequencing data and is a multi-stage bioinformatics process yielding sets of potential circRNA candidates that require further analyses. While a number of tools for the prediction process already exist, publicly available analysis tools for further characterization are rare. Our work provides researchers with a harmonized workflow that covers different stages of in silico circRNA analyses, from prediction to first functional insights. RESULTS: Here, we present circtools, a modular, Python-based framework for computational circRNA analyses. The software includes modules for circRNA detection, internal sequence reconstruction, quality checking, statistical testing, screening for enrichment of RBP binding sites, differential exon RNase R resistance and circRNA-specific primer design. circtools supports researchers with visualization options and data export into commonly used formats. AVAILABILITY AND IMPLEMENTATION: circtools is available via https://github.com/dieterich-lab/circtools and http://circ.tools under GPLv3.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tobias Jakobi, Alexey Uvarovskii, Christoph Dieterich |
Bioinform. | 3 |
| 2019 | On the optimal design of metabolic RNA labeling experimentsabstractMassively parallel RNA sequencing (RNA-seq) in combination with metabolic labeling has become the de facto standard approach to study alterations in RNA transcription, processing or decay. Regardless of advances in the experimental protocols and techniques, every experimentalist needs to specify the key aspects of experimental design: For example, which protocol should be used (biochemical separation vs. nucleotide conversion) and what is the optimal labeling time? In this work, we provide approximate answers to these questions using the asymptotic theory of optimal design. Specifically, we investigate, how the variance of degradation rate estimates depends on the time and derive the optimal time for any given degradation rate. Subsequently, we show that an increase in sample numbers should be preferred over an increase in sequencing depth. Lastly, we provide some guidance on use cases when laborious biochemical separation outcompetes recent nucleotide conversion based methods (such as SLAMseq) and show, how inefficient conversion influences the precision of estimates. Code and documentation can be found at https://github.com/dieterich-lab/DesignMetabolicRNAlabeling. Alexey Uvarovskii, Isabel S. Naarmann-de Vries, Christoph Dieterich |
PLoS Comput. Biol. | 3 |
| 2017 | Flexbar 3.0 - SIMD and multicore parallelizationabstractMOTIVATION: High-throughput sequencing machines can process many samples in a single run. For Illumina systems, sequencing reads are barcoded with an additional DNA tag that is contained in the respective sequencing adapters. The recognition of barcode and adapter sequences is hence commonly needed for the analysis of next-generation sequencing data. Flexbar performs demultiplexing based on barcodes and adapter trimming for such data. The massive amounts of data generated on modern sequencing machines demand that this preprocessing is done as efficiently as possible. RESULTS: We present Flexbar 3.0, the successor of the popular program Flexbar. It employs now twofold parallelism: multi-threading and additionally SIMD vectorization. Both types of parallelism are used to speed-up the computation of pair-wise sequence alignments, which are used for the detection of barcodes and adapters. Furthermore, new features were included to cover a wide range of applications. We evaluated the performance of Flexbar based on a simulated sequencing dataset. Our program outcompetes other tools in terms of speed and is among the best tools in the presented quality benchmark. AVAILABILITY AND IMPLEMENTATION: https://github.com/seqan/flexbar. CONTACT: [email protected] or [email protected]. Johannes T. Roehr, Christoph Dieterich, Knut Reinert |
Bioinform. | 2 |
| 2017 | pulseR: Versatile computational analysis of RNA turnover from metabolic labeling experimentsabstractMOTIVATION: Metabolic labelling of RNA is a well-established and powerful method to estimate RNA synthesis and decay rates. The pulseR R package simplifies the analysis of RNA-seq count data that emerge from corresponding pulse-chase experiments. RESULTS: The pulseR package provides a flexible interface and readily accommodates numerous different experimental designs. To our knowledge, it is the first publicly available software solution that models count data with the more appropriate negative-binomial model. Moreover, pulseR handles labelled and unlabelled spike-in sets in its workflow and accounts for potential labeling biases (e.g. number of uridine residues). AVAILABILITY AND IMPLEMENTATION: The pulseR package is freely available at https://github.com/dieterich-lab/pulseR under the GPLv3.0 licence. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexey Uvarovskii, Christoph Dieterich |
Bioinform. | 2 |
| 2017 | JACUSA: site-specific identification of RNA editing events from replicate sequencing dataabstractBACKGROUND: RNA editing is a co-transcriptional modification that increases the molecular diversity, alters secondary structure and protein coding sequences by changing the sequence of transcripts. The most common RNA editing modification is the single base substitution (A→I) that is catalyzed by the members of the Adenosine deaminases that act on RNA (ADAR) family. Typically, editing sites are identified as RNA-DNA-differences (RDDs) in a comparison of genome and transcriptome data from next-generation sequencing experiments. However, a method for robust detection of site-specific editing events from replicate RNA-seq data has not been published so far. Even more surprising, condition-specific editing events, which would show up as differences in RNA-RNA comparisons (RRDs) and depend on particular cellular states, are rarely discussed in the literature. RESULTS: We present JACUSA, a versatile one-stop solution to detect single nucleotide variant positions from comparing RNA-DNA and/or RNA-RNA sequencing samples. The performance of JACUSA has been carefully evaluated and compared to other variant callers in an in silico benchmark. JACUSA outperforms other algorithms in terms of the F measure, which combines precision and recall, in all benchmark scenarios. This performance margin is highest for the RNA-RNA comparison scenario. We further validated JACUSA's performance by testing its ability to detect A→I events using sequencing data from a human cell culture experiment and publicly available RNA-seq data from Drosophila melanogaster heads. To this end, we performed whole genome and RNA sequencing of HEK-293 cells on samples with lowered activity of candidate RNA editing enzymes. JACUSA has a higher recall and comparable precision for detecting true editing sites in RDD comparisons of HEK-293 data. Intriguingly, JACUSA captures most A→I events from RRD comparisons of RNA sequencing data derived from Drosophila and HEK-293 data sets. CONCLUSION: Our software JACUSA detects single nucleotide variants by comparing data from next-generation sequencing experiments (RNA-DNA or RNA-RNA). In practice, JACUSA shows higher recall and comparable precision in detecting A→I sites from RNA-DNA comparisons, while showing higher precision and recall in RNA-RNA comparisons. Michael Piechotta, Emanuel Wyler, Uwe Ohler, Markus Landthaler, Christoph Dieterich |
BMC Bioinform. | 5 |
| 2016 | Specific identification and quantification of circular RNAs from sequencing dataabstractMOTIVATION: Circular RNAs (circRNAs) are a poorly characterized class of molecules that have been identified decades ago. Emerging high-throughput sequencing methods as well as first reports on confirmed functions have sparked new interest in this RNA species. However, the computational detection and quantification tools are still limited. RESULTS: We developed the software tandem, DCC and CircTest DCC uses output from the STAR read mapper to systematically detect back-splice junctions in next-generation sequencing data. DCC applies a series of filters and integrates data across replicate sets to arrive at a precise list of circRNA candidates. We assessed the detection performance of DCC on a newly generated mouse brain data set and publicly available sequencing data. Our software achieves a much higher precision than state-of-the-art competitors at similar sensitivity levels. Moreover, DCC estimates circRNA versus host gene expression from counting junction and non-junction reads. These read counts are finally used to test for host gene-independence of circRNA expression across different experimental conditions by our R package CircTest We demonstrate the benefits of this approach on previously reported age-dependent circRNAs in the fruit fly. AVAILABILITY AND IMPLEMENTATION: The source code of DCC and CircTest is licensed under the GNU General Public Licence (GPL) version 3 and available from https://github.com/dieterich-lab/[DCC or CircTest]. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Cheng 0007, Franziska Metge, Christoph Dieterich |
Bioinform. | 3 |
| 2014 | Detection of active transcription factor binding sites with the combination of DNase hypersensitivity and histone modificationsabstractMOTIVATION: The identification of active transcriptional regulatory elements is crucial to understand regulatory networks driving cellular processes such as cell development and the onset of diseases. It has recently been shown that chromatin structure information, such as DNase I hypersensitivity (DHS) or histone modifications, significantly improves cell-specific predictions of transcription factor binding sites. However, no method has so far successfully combined both DHS and histone modification data to perform active binding site prediction. RESULTS: We propose here a method based on hidden Markov models to integrate DHS and histone modifications occupancy for the detection of open chromatin regions and active binding sites. We have created a framework that includes treatment of genomic signals, model training and genome-wide application. In a comparative analysis, our method obtained a good trade-off between sensitivity versus specificity and superior area under the curve statistics than competing methods. Moreover, our technique does not require further training or sequence information to generate binding location predictions. Therefore, the method can be easily applied on new cell types and allow flexible downstream analysis such as de novo motif finding. AVAILABILITY AND IMPLEMENTATION: Our framework is available as part of the Regulatory Genomics Toolbox. The software information and all benchmarking data are available at http://costalab.org/wp/dh-hmm. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Eduardo G. Gusmão, Christoph Dieterich, Martin Zenke, Ivan G. Costa |
Bioinform. | 2 |
| 2013 | ACCUSA2: multi-purpose SNV calling enhanced by probabilistic integration of quality scoresabstractSUMMARY: Direct comparisons of assembled short-read stacks are one way to identify single-nucleotide variants. Single-nucleotide variant detection is especially challenging across samples with different read depths (e.g. RNA-Seq) and high-background levels (e.g. selection experiments). We present ACCUSA2 to identify variant positions where nucleotide frequency spectra differ between two samples. To this end, ACCUSA2 integrates quality scores for base calling and read mapping into a common framework. Our benchmarks demonstrate that ACCUSA2 is superior to a state-of-the-art SNV caller in situations of diverging read depths and reliably detects subtle differences among sample nucleotide frequency spectra. Additionally, we show that ACCUSA2 is fast and robust against base quality score deviations. AVAILABILITY: ACCUSA2 is available free of charge to academic users and may be obtained from https://bbc.mdc-berlin.de/software. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael Piechotta, Christoph Dieterich |
Bioinform. | 2 |
| 2010 | ACCUSA - accurate SNP calling on draft genomesabstractSUMMARY: Next generation sequencing technologies facilitate genome-wide analysis of several biological processes. We are interested in whole-genome genotyping. To our knowledge, none of the existing single nucleotide polymorphism (SNP) callers consider the quality of the reference genome, which is not necessary for high-quality assemblies of well-studied model organisms. However, most genome projects will remain in draft status with little to no genome assembly improvement due to time and financial constraints. Here, we present a simple yet elegant solution ('ACCUSA') that considers both the read qualities as well as the reference genome's quality using a Bayesian framework. We demonstrate that ACCUSA is as good as the current SNP calling software in detecting true SNPs. More importantly, ACCUSA does not call spurious SNPs, which originate from a poor reference sequence. AVAILABILITY: ACCUSA is available free of charge to academic users and may be obtained from ftp://bbc.mdc-berlin.de/software. ACCUSA is programmed in JAVA 6 and runs on any platform with JAVA support. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sebastian Fröhler, Christoph Dieterich |
Bioinform. | 2 |
| 2005 | SITEBLAST-rapid and sensitive local alignment of genomic sequences employing motif anchorsabstractMOTIVATION: Comparative sequence analysis is the essence of many approaches to genome annotation. Heuristic alignment algorithms utilize similar seed pairs to anchor an alignment. Some applications of local alignment algorithms (e.g. phylogenetic footprinting) would benefit from including prior knowledge (e.g. binding site motifs) in the alignment building process. RESULTS: We introduce predefined sequence patterns as anchor points into a heuristic local alignment strategy. We extended the BLASTZ program for this purpose. A set of seed patterns is either given as consensus sequences in IUPAC code or position-weight-matrices. Phylogenetic footprinting of promoter regions is one of many potential applications for the SITEBLAST software. AVAILABILITY: The source code is freely available to the academic community from http://corg.molgen.mpg.de/software Morris Michael, Christoph Dieterich, Martin Vingron |
Bioinform. | 2 |
| 2005 | Ab initio identification of putative human transcription factor binding sites by comparative genomicsabstractBACKGROUND: Understanding transcriptional regulation of gene expression is one of the greatest challenges of modern molecular biology. A central role in this mechanism is played by transcription factors, which typically bind to specific, short DNA sequence motifs usually located in the upstream region of the regulated genes. We discuss here a simple and powerful approach for the ab initio identification of these cis-regulatory motifs. The method we present integrates several elements: human-mouse comparison, statistical analysis of genomic sequences and the concept of coregulation. We apply it to a complete scan of the human genome. RESULTS: By using the catalogue of conserved upstream sequences collected in the CORG database we construct sets of genes sharing the same overrepresented motif (short DNA sequence) in their upstream regions both in human and in mouse. We perform this construction for all possible motifs from 5 to 8 nucleotides in length and then filter the resulting sets looking for two types of evidence of coregulation: first, we analyze the Gene Ontology annotation of the genes in the set, searching for statistically significant common annotations; second, we analyze the expression profiles of the genes in the set as measured by microarray experiments, searching for evidence of coexpression. The sets which pass one or both filters are conjectured to contain a significant fraction of coregulated genes, and the upstream motifs characterizing the sets are thus good candidates to be the binding sites of the TF's involved in such regulation. In this way we find various known motifs and also some new candidate binding sites. CONCLUSION: We have discussed a new integrated algorithm for the "ab initio" identification of transcription factor binding sites in the human genome. The method is based on three ingredients: comparative genomics, overrepresentation, different types of coregulation. The method is applied to a full-scan of the human genome, giving satisfactory results. Davide Corà, Carl Herrmann, Christoph Dieterich, Ferdinando Di Cunto, Paolo Provero, Michele Caselle |
BMC Bioinform. | 3 |
| 2004 | Suboptimal Local Alignments Across Multiple Scoring Schemes
Morris Michael, Christoph Dieterich, Jens Stoye |
WABI | 2 |