EDBT 2026 Demo / reviewers in the wild / expert
Manja Marz
dblp:03/1787
· DBLP profile ↗
16ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-4783-8823ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | diffMONT : predicting methylation-specific PCR biomarkers based on nanopore sequencing data for clinical applicationabstractMOTIVATION: DNA methylation serves as a key biomarker in clinical diagnostics, especially in cancer detection. With methylation-specific PCR (MSP), a widely used approach, patient samples can be screened fast and efficiently for differential methylation. During MSP, methylated regions are selectively amplified with specific primers. With nanopore sequencing, knowledge about DNA methylation is generated during direct DNA sequencing without needing pretreatment of the DNA. Multiple methods, mainly developed for whole-genome bisulfite sequencing (WGBS) data, exist to predict differentially methylated regions (DMRs) in the genome. However, the predicted DMRs are often very large and not sufficiently discriminating to generate meaningful results in MSP, creating a gap between theoretical cancer marker research and practical application, as no tool currently provides methylation difference predictions tailored for PCR-based diagnostics. RESULTS: Here, we present diffMONT, a tool that predicts differentially methylated regions specifically suited for MSP primer design, enabling rapid translation into practical applications. diffMONT takes into account (i) the specific length of primer and amplicon regions, (ii) the fact that one condition should be unmethylated, and (iii) a minimal required amount of differentially methylated cytosines within the primer regions. We compared the results of diffMONT to metilene and DSS based on a publicly available nanopore sequencing dataset and show that the regions predicted by diffMONT are more specific toward hypermethylated regions. diffMONT accelerates the design of methylation-specific diagnostic assays, bridging the gap between theoretical research and clinical application. AVAILABILITY AND IMPLEMENTATION: The source code for diffMONT, an open-source Python-based tool, is available at https://github.com/rnajena/diffMONT/, with an archived release under https://zenodo.org/records/17641031. Daria Meyer, Emanuel Barth, Laura Wiehle, Manja Marz |
Bioinform. | 4 |
| 2025 | Go Game Capture and Reconstruction of Missing Moves Using Deep Learning Techniques
Roman Gerloff, Manja Marz, Florian Wahl |
CoG | 2 |
| 2023 | VIRify: An integrated detection, annotation and taxonomic classification pipeline using virus-specific protein profile hidden Markov modelsabstractThe study of viral communities has revealed the enormous diversity and impact these biological entities have on various ecosystems. These observations have sparked widespread interest in developing computational strategies that support the comprehensive characterisation of viral communities based on sequencing data. Here we introduce VIRify, a new computational pipeline designed to provide a user-friendly and accurate functional and taxonomic characterisation of viral communities. VIRify identifies viral contigs and prophages from metagenomic assemblies and annotates them using a collection of viral profile hidden Markov models (HMMs). These include our manually-curated profile HMMs, which serve as specific taxonomic markers for a wide range of prokaryotic and eukaryotic viral taxa and are thus used to reliably classify viral contigs. We tested VIRify on assemblies from two microbial mock communities, a large metagenomics study, and a collection of publicly available viral genomic sequences from the human gut. The results showed that VIRify could identify sequences from both prokaryotic and eukaryotic viruses, and provided taxonomic classifications from the genus to the family rank with an average accuracy of 86.6%. In addition, VIRify allowed the detection and taxonomic classification of a range of prokaryotic and eukaryotic viruses present in 243 marine metagenomic assemblies. Finally, the use of VIRify led to a large expansion in the number of taxonomically classified human gut viral sequences and the improvement of outdated and shallow taxonomic classifications. Overall, we demonstrate that VIRify is a novel and powerful resource that offers an enhanced capability to detect a broad range of viral contigs and taxonomically classify them. Guillermo Rangel-Pineros, Alexandre Almeida, Martin Beracochea, Ekaterina A. Sakharova, Manja Marz, Martin Hölzer, Robert D. Finn |
PLoS Comput. Biol. | 5 |
| 2021 | Computational strategies to combat COVID-19: useful tools to accelerate SARS-CoV-2 and coronavirus researchabstractSARS-CoV-2 (severe acute respiratory syndrome coronavirus 2) is a novel virus of the family Coronaviridae. The virus causes the infectious disease COVID-19. The biology of coronaviruses has been studied for many years. However, bioinformatics tools designed explicitly for SARS-CoV-2 have only recently been developed as a rapid reaction to the need for fast detection, understanding and treatment of COVID-19. To control the ongoing COVID-19 pandemic, it is of utmost importance to get insight into the evolution and pathogenesis of the virus. In this review, we cover bioinformatics workflows and tools for the routine detection of SARS-CoV-2 infection, the reliable analysis of sequencing data, the tracking of the COVID-19 pandemic and evaluation of containment measures, the study of coronavirus evolution, the discovery of potential drug targets and development of therapeutic strategies. For each tool, we briefly describe its use case and how it advances research specifically for SARS-CoV-2. All tools are free to use and available online, either through web applications or public code repositories. Contact:[email protected]. Franziska Hufsky, Kevin Lamkiewicz, Alexandre Almeida, Abdel Aouacheria, Cecilia N. Arighi, Alex Bateman, Jan Baumbach, Niko Beerenwinkel, Christian Brandt, Marco Cacciabue, Sara Chuguransky, Oliver Drechsel, Robert D. Finn, Adrian Fritz, Stephan Fuchs, Georges Hattab, Anne-Christin Hauschild, Dominik Heider, Marie Hoffmann, Martin Hölzer, Stefan Hoops, Lars Kaderali, Ioanna Kalvari, Max von Kleist, Renó Kmiecinski, Denise Kühnert, Gorka Lasso, Pieter Libin, Markus List, Hannah F. Löchel, Maria Jesus Martin, Roman Martin, Julian O. Matschinske, Alice C. McHardy, Pedro Mendes 0001, Jaina Mistry, Vincent Navratil, Eric P. Nawrocki, Áine Niamh O'toole, Nancy Ontiveros-Palacios, Anton I. Petrov, Guillermo Rangel-Pineros, Nicole Redaschi, Susanne Reimering, Knut Reinert, Lorna J. Richardson, David L. Robertson, Sepideh Sadegh, Joshua B. Singer, Kristof Theys, Chris Upton, Marius Welzel, Lowri Williams, Manja Marz |
Briefings Bioinform. | 55 |
| 2021 | EpiDope: a deep neural network for linear B-cell epitope predictionabstractMOTIVATION: By binding to specific structures on antigenic proteins, the so-called epitopes, B-cell antibodies can neutralize pathogens. The identification of B-cell epitopes is of great value for the development of specific serodiagnostic assays and the optimization of medical therapy. However, identifying diagnostically or therapeutically relevant epitopes is a challenging task that usually involves extensive laboratory work. In this study, we show that the time, cost and labor-intensive process of epitope detection in the lab can be significantly reduced using in silico prediction. RESULTS: Here, we present EpiDope, a python tool which uses a deep neural network to detect linear B-cell epitope regions on individual protein sequences. With an area under the curve between 0.67 ± 0.07 in the receiver operating characteristic curve, EpiDope exceeds all other currently used linear B-cell epitope prediction tools. Our software is shown to reliably predict linear B-cell epitopes of a given protein sequence, thus contributing to a significant reduction of laboratory experiments and costs required for the conventional approach. AVAILABILITYAND IMPLEMENTATION: EpiDope is available on GitHub (http://github.com/mcollatz/EpiDope). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Maximilian Collatz, Florian Mock, Emanuel Barth, Martin Hölzer, Konrad Sachse, Manja Marz |
Bioinform. | 6 |
| 2021 | EpiDope: a deep neural network for linear B-cell epitope predictionabstractBioinformatics (2021) doi: 10.1093/bioinformatics/btaa773 In the originally published version of this manuscript, there was an erroneous omission in the Funding section. The section should read: “This work was funded in the framework of the national research network InfectControl, project "Molecular serology for rapid determination of vaccination titers (STIKO Serology)", which was financially supported by the Federal Ministry of Education and Research (BMBF) of Germany under grant 03ZZ0820A. This work was further supported by the Collaborative Research Center/Transregio 124 (FungiNet; number 210879364), project B5, funded by Deutsche Forschungsgemeinschaft (DFG). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.” instead of: “This work was funded in the framework of the national research network InfectControl, project ‘Molecular serology for rapid determination of vaccination titers (STIKO Serology)’, which was financially supported by the Federal Ministry of Education and Research (BMBF) of Germany [03ZZ0820A]. The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.” This error has now been corrected online. Maximilian Collatz, Florian Mock, Emanuel Barth, Martin Hölzer, Konrad Sachse, Manja Marz |
Bioinform. | 6 |
| 2021 | PoSeiDon: a Nextflow pipeline for the detection of evolutionary recombination events and positive selectionabstractSUMMARY: PoSeiDon is an easy-to-use pipeline that helps researchers to find recombination events and sites under positive selection in protein-coding sequences. By entering homologous sequences, PoSeiDon builds an alignment, estimates a best-fitting substitution model and performs a recombination analysis followed by the construction of all corresponding phylogenies. Finally, significantly positive selected sites are detected according to different models for the full alignment and possible recombination fragments. The results of PoSeiDon are summarized in a user-friendly HTML page providing all intermediate results and the graphical representation of recombination events and positively selected sites. AVAILABILITY AND IMPLEMENTATION: PoSeiDon is freely available at https://github.com/hoelzer/poseidon. The pipeline is implemented in Nextflow with Docker support and processes the output of various tools. Martin Hölzer, Manja Marz |
Bioinform. | 2 |
| 2021 | VIDHOP, viral host prediction with deep learningabstractMOTIVATION: Zoonosis, the natural transmission of infections from animals to humans, is a far-reaching global problem. The recent outbreaks of Zikavirus, Ebolavirus and Coronavirus are examples of viral zoonosis, which occur more frequently due to globalization. In case of a virus outbreak, it is helpful to know which host organism was the original carrier of the virus to prevent further spreading of viral infection. Recent approaches aim to predict a viral host based on the viral genome, often in combination with the potential host genome and arbitrarily selected features. These methods are limited in the number of different hosts they can predict or the accuracy of the prediction. RESULTS: Here, we present a fast and accurate deep learning approach for viral host prediction, which is based on the viral genome sequence only. We tested our deep neural network (DNN) on three different virus species (influenza A virus, rabies lyssavirus and rotavirus A). We achieved for each virus species an AUC between 0.93 and 0.98, allowing highly accurate predictions while using only fractions (100-400 bp) of the viral genome sequences. We show that deep neural networks are suitable to predict the host of a virus, even with a limited amount of sequences and highly unbalanced available data. The trained DNNs are the core of our virus-host prediction tool VIrus Deep learning HOst Prediction (VIDHOP). VIDHOP also allows the user to train and use models for other viruses. AVAILABILITY AND IMPLEMENTATION: VIDHOP is freely available under https://github.com/flomock/vidhop. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Florian Mock, Adrian Viehweger, Emanuel Barth, Manja Marz |
Bioinform. | 4 |
| 2019 | Global importance of RNA secondary structures in protein-coding sequencesabstractMOTIVATION: The protein-coding sequences of messenger RNAs are the linear template for translation of the gene sequence into protein. Nevertheless, the RNA can also form secondary structures by intramolecular base-pairing. RESULTS: We show that the nucleotide distribution within codons is biased in all taxa of life on a global scale. Thereby, RNA secondary structures that require base-pairing between the position 1 of a codon with the position 1 of an opposing codon (here named RNA secondary structure class c1) are under-represented. We conclude that this bias may result from the co-evolution of codon sequence and mRNA secondary structure, suggesting that RNA secondary structures are generally important in protein-coding regions of mRNAs. The above result also implies that codon position 2 has a smaller influence on the amino acid choice than codon position 1. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Markus Fricke, Ruman Gerst, Bashar Ibrahim, Michael Niepmann, Manja Marz |
Bioinform. | 5 |
| 2016 | Prediction of conserved long-range RNA-RNA interactions in full viral genomesabstractMOTIVATION: Long-range RNA-RNA interactions (LRIs) play an important role in viral replication, however, only a few of these interactions are known and only for a small number of viral species. Up to now, it has been impossible to screen a full viral genome for LRIs experimentally or in silico Most known LRIs are cross-reacting structures (pseudoknots) undetectable by most bioinformatical tools. RESULTS: We present LRIscan, a tool for the LRI prediction in full viral genomes based on a multiple genome alignment. We confirmed 14 out of 16 experimentally known and evolutionary conserved LRIs in genome alignments of HCV, Tombusviruses, Flaviviruses and HIV-1. We provide several promising new interactions, which include compensatory mutations and are highly conserved in all considered viral sequences. Furthermore, we provide reactivity plots highlighting the hot spots of predicted LRIs. AVAILABILITY AND IMPLEMENTATION: Source code and binaries of LRIscan freely available for download at http://www.rna.uni-jena.de/en/supplements/lriscan/, implemented in Ruby/C ++ and supported on Linux and Windows. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Markus Fricke, Manja Marz |
Bioinform. | 2 |
| 2014 | Explorative Analysis of Heterogeneous, Unstructured, and Uncertain Data - A Computer Science Perspective on Biodiversity ResearchabstractWe outline a blueprint for the development of new computer science approaches for the management and analysis of big data problems for biodiversity science. Such problems are
characterized by a combination of different data sources each of which owns at least one of the typical characteristics of big data (volume, variety, velocity, or veracity). For these problems, we envision a solution that covers different aspects of integrating data sources and algorithms for their analysis on one of the following three layers: At the data layer, there are various data archives of heterogeneous, unstructured, and uncertain data. At the functional layer, the data are analyzed for each archive individually. At the meta-layer, multiple functional archives are combined for complex analysis. Clemens Beckstein, Sebastian Böcker, Martin Bogdan, Helge Bruelheide, H. Martin Bücker, Joachim Denzler, Peter Dittrich, Ivo Grosse, Alexander Hinneburg, Birgitta König-Ries, Felicitas Löffler, Manja Marz, Matthias Müller-Hannemann, Wolf Zimmermann |
DATA | 12 |
| 2014 | Challenges in RNA virus bioinformaticsabstractMOTIVATION: Computer-assisted studies of structure, function and evolution of viruses remains a neglected area of research. The attention of bioinformaticians to this interesting and challenging field is far from commensurate with its medical and biotechnological importance. It is telling that out of >200 talks held at ISMB 2013, the largest international bioinformatics conference, only one presentation explicitly dealt with viruses. In contrast to many broad, established and well-organized bioinformatics communities (e.g. structural genomics, ontologies, next-generation sequencing, expression analysis), research groups focusing on viruses can probably be counted on the fingers of two hands. RESULTS: The purpose of this review is to increase awareness among bioinformatics researchers about the pressing needs and unsolved problems of computational virology. We focus primarily on RNA viruses that pose problems to many standard bioinformatics analyses owing to their compact genome organization, fast mutation rate and low evolutionary conservation. We provide an overview of tools and algorithms for handling viral sequencing data, detecting functionally important RNA structures, classifying viral proteins into families and investigating the origin and evolution of viruses. Manja Marz, Niko Beerenwinkel, Christian Drosten, Markus Fricke, Dmitrij Frishman, Ivo L. Hofacker, Dieter Hoffmann, Martin Middendorf, Thomas Rattei, Peter F. Stadler, Armin Töpfer |
Bioinform. | 1 |
| 2013 | POMAGO: Multiple Genome-Wide Alignment Tool for Bacteria
Nicolas Wieseke, Marcus Lechner, Marcus Ludwig, Manja Marz |
ISBRA | 4 |
| 2013 | Distribution of Graph-Distances in Boltzmann Ensembles of RNA Secondary Structures
Rolf Backofen, Markus Fricke, Manja Marz, Jing Qin 0006, Peter F. Stadler |
WABI | 3 |
| 2011 | RNA-RNA interaction prediction based on multiple sequence alignmentsabstractMOTIVATION: Many computerized methods for RNA-RNA interaction structure prediction have been developed. Recently, O(N(6)) time and O(N(4)) space dynamic programming algorithms have become available that compute the partition function of RNA-RNA interaction complexes. However, few of these methods incorporate the knowledge concerning related sequences, thus relevant evolutionary information is often neglected from the structure determination. Therefore, it is of considerable practical interest to introduce a method taking into consideration both: thermodynamic stability as well as sequence/structure covariation. RESULTS: We present the a priori folding algorithm ripalign, whose input consists of two (given) multiple sequence alignments (MSA). ripalign outputs (i) the partition function, (ii) base pairing probabilities, (iii) hybrid probabilities and (iv) a set of Boltzmann-sampled suboptimal structures consisting of canonical joint structures that are compatible to the alignments. Compared to the single sequence-pair folding algorithm rip, ripalign requires negligible additional memory resource but offers much better sensitivity and specificity, once alignments of suitable quality are given. ripalign additionally allows to incorporate structure constraints as input parameters. AVAILABILITY: The algorithm described here is implemented in C as part of the rip package. Andrew X. Li, Manja Marz, Jing Qin 0006, Christian M. Reidys |
Bioinform. | 2 |
| 2011 | Proteinortho: Detection of (Co-)Orthologs in Large-Scale AnalysisabstractBACKGROUND: Orthology analysis is an important part of data analysis in many areas of bioinformatics such as comparative genomics and molecular phylogenetics. The ever-increasing flood of sequence data, and hence the rapidly increasing number of genomes that can be compared simultaneously, calls for efficient software tools as brute-force approaches with quadratic memory requirements become infeasible in practise. The rapid pace at which new data become available, furthermore, makes it desirable to compute genome-wide orthology relations for a given dataset rather than relying on relations listed in databases. RESULTS: The program Proteinortho described here is a stand-alone tool that is geared towards large datasets and makes use of distributed computing techniques when run on multi-core hardware. It implements an extended version of the reciprocal best alignment heuristic. We apply Proteinortho to compute orthologous proteins in the complete set of all 717 eubacterial genomes available at NCBI at the beginning of 2009. We identified thirty proteins present in 99% of all bacterial proteomes. CONCLUSIONS: Proteinortho significantly reduces the required amount of memory for orthology analysis compared to existing tools, allowing such computations to be performed on off-the-shelf hardware. Marcus Lechner, Sven Findeiß, Lydia Steiner, Manja Marz, Peter F. Stadler, Sonja J. Prohaska |
BMC Bioinform. | 4 |