VLDB 2026 Research / reviewers in the wild / expert
Manuel Holtgrewe
dblp:64/7481
· DBLP profile ↗
10ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0002-3051-1763ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection |
0.6 | 1 | 2022 | ClearCNV: CNV calling from NGS panel data in the presence of ambiguity and noise · Bioinform. 2022 |
Bioinformatics and computational biology › genomics
variant calling |
0.6 | 1 | 2022 | ClearCNV: CNV calling from NGS panel data in the presence of ambiguity and noise · Bioinform. 2022 |
Bioinformatics and computational biology › bioinformatics infrastructure
sequencing data management |
0.4 | 1 | 2020 | DigestiFlow: from BCL to FASTQ with ease · Bioinform. 2020 |
Bioinformatics and computational biology › sequence analysis
sequencing data processing |
0.4 | 1 | 2020 | DigestiFlow: from BCL to FASTQ with ease · Bioinform. 2020 |
Bioinformatics and computational biology
genomics |
0.3 | 1 | 2017 | HLA-MA: simple yet powerful matching of samples using HLA typing results · Bioinform. 2017 |
Bioinformatics and computational biology › genomics
high-throughput sequencing |
0.3 | 1 | 2017 | HLA-MA: simple yet powerful matching of samples using HLA typing results · Bioinform. 2017 |
Bioinformatics and computational biology › immunoinformatics
HLA typing |
0.3 | 1 | 2017 | HLA-MA: simple yet powerful matching of samples using HLA typing results · Bioinform. 2017 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly |
0.2 | 1 | 2015 | Methods for the detection and assembly of novel sequence in high-throughput sequencing data · Bioinform. 2015 |
Bioinformatics and computational biology › genomics › structural variation
structural variant detection |
0.2 | 1 | 2015 | Methods for the detection and assembly of novel sequence in high-throughput sequencing data · Bioinform. 2015 |
Bioinformatics and computational biology › sequence analysis
sequencing error correction |
0.2 | 1 | 2014 | Fiona: a parallel and automatic strategy for read error correction · Bioinform. 2014 |
Bioinformatics and computational biology › sequence analysis
read mapping |
0.1 | 1 | 2012 | RazerS 3: Faster, fully sensitive read mapping · Bioinform. 2012 |
Parallel and multicore computing › parallel computing
multicore parallelism |
0.1 | 1 | 2014 | Fiona: a parallel and automatic strategy for read error correction · Bioinform. 2014 |
Parallel and multicore computing
parallel computing |
0.1 | 1 | 2014 | Fiona: a parallel and automatic strategy for read error correction · Bioinform. 2014 |
Algorithms and data structures › sequence algorithms › string algorithms
sequence alignment |
0.0 | 1 | 2012 | RazerS 3: Faster, fully sensitive read mapping · Bioinform. 2012 |
Algorithms and data structures › sequence algorithms
string algorithms |
0.0 | 1 | 2012 | RazerS 3: Faster, fully sensitive read mapping · Bioinform. 2012 |
Methods — techniques the papers use, named apart from their topics
homogeneous subset identification · 0.6enrichment kit assignment · 0.6workflow automation · 0.4statistical error model · 0.4partial suffix array · 0.4parameter estimation · 0.4q-gram counting · 0.3paired-end read analysis · 0.2de novo assembly · 0.2seed-based filter · 0.1myers' bitvector algorithm · 0.1myers' bit-vector algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | ClearCNV: CNV calling from NGS panel data in the presence of ambiguity and noiseabstractMOTIVATION: While the identification of small variants in panel sequencing data can be considered a solved problem, the identification of larger, multi-exon copy number variants (CNVs) still poses a considerable challenge. Thus, CNV calling has not been established in all laboratories performing panel sequencing. At the same time, such laboratories have accumulated large datasets and thus have the need to identify CNVs on their data to close the diagnostic gap. RESULTS: In this article, we present our method clearCNV that addresses this need in two ways. First, it helps laboratories to properly assign datasets to enrichment kits. Based on homogeneous subsets of data, clearCNV identifies CNVs affecting the targeted regions. Using real-world datasets and validation, we show that our method is highly competitive with previous methods and preferable in terms of specificity. AVAILABILITY AND IMPLEMENTATION: The software is available for free under a permissible license at https://github.com/bihealth/clear-cnv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vinzenz May, Leonard Koch, Björn Fischer-Zirnsak, Denise Horn, Petra Gehle, Uwe Kornak, Dieter Beule, Manuel Holtgrewe |
Bioinform. | 8 |
| 2020 | DigestiFlow: from BCL to FASTQ with easeabstractAbstract Summary Management of raw-sequencing data and its pre-processing (conversion into sequences and demultiplexing) remains a challenging topic for groups running sequencing devices. They face many challenges in such efforts and solutions ranging from manual management of spreadsheets to very complex and customized laboratory information management systems handling much more than just sequencing raw data. In this article, we describe the software package DigestiFlow that focuses on the management of Illumina flow cell sample sheets and raw data. It allows for automated extraction of information from flow cell data and management of sample sheets. Furthermore, it allows for the automated and reproducible conversion of Illumina base calls to sequences and the demultiplexing thereof using bcl2fastq and Picard Tools, followed by quality control report generation. Availability and implementation The software is available under the MIT license at https://github.com/bihealth/digestiflow-server. The client software components are available via Bioconda. Supplementary information Supplementary data are available at Bioinformatics online. Manuel Holtgrewe, Clemens Messerschmidt, Mikko Nieminen, Dieter Beule |
Bioinform. | 1 |
| 2017 | HLA-MA: simple yet powerful matching of samples using HLA typing resultsabstractSUMMARY: We propose the simple method HLA-MA for consistency checking in pipelines operating on human HTS data. The method is based on the HLA typing result of the state-of-the-art method OptiType. Provided that there is sufficient coverage of the HLA loci, comparing HLA types allows for simple, fast and robust matching of samples from whole genome, exome and RNA-seq data. Our approach uses information from small but genetically highly variable regions and thus complements approaches that rely on genome or exon-wide variant profiles. AVAILABILITY AND IMPLEMENTATION: The software is implemented In Python 3 and freely available under the MIT license at https://github.com/bihealth/hlama and via Bioconda. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Clemens Messerschmidt, Manuel Holtgrewe, Dieter Beule |
Bioinform. | 2 |
| 2015 | Methods for the detection and assembly of novel sequence in high-throughput sequencing dataabstractMOTIVATION: Large insertions of novel sequence are an important type of structural variants. Previous studies used traditional de novo assemblers for assembling non-mapping high-throughput sequencing (HTS) or capillary reads and then tried to anchor them in the reference using paired read information. RESULTS: We present approaches for detecting insertion breakpoints and targeted assembly of large insertions from HTS paired data: BASIL and ANISE. On near identity repeats that are hard for assemblers, ANISE employs a repeat resolution step. This results in far better reconstructions than obtained by the compared methods. On simulated data, we found our insert assembler to be competitive with the de novo assemblers ABYSS and SGA while yielding already anchored inserted sequence as opposed to unanchored contigs as from ABYSS/SGA. On real-world data, we detected novel sequence in a human individual and thoroughly validated the assembled sequence. ANISE was found to be superior to the competing tool MindTheGap on both simulated and real-world data. AVAILABILITY AND IMPLEMENTATION: ANISE and BASIL are available for download at http://www.seqan.de/projects/herbarium under a permissive open source license. Manuel Holtgrewe, Léon Kuchenbecker, Knut Reinert |
Bioinform. | 1 |
| 2014 | Fiona: a parallel and automatic strategy for read error correctionabstractMOTIVATION: Automatic error correction of high-throughput sequencing data can have a dramatic impact on the amount of usable base pairs and their quality. It has been shown that the performance of tasks such as de novo genome assembly and SNP calling can be dramatically improved after read error correction. While a large number of methods specialized for correcting substitution errors as found in Illumina data exist, few methods for the correction of indel errors, common to technologies like 454 or Ion Torrent, have been proposed. RESULTS: We present Fiona, a new stand-alone read error-correction method. Fiona provides a new statistical approach for sequencing error detection and optimal error correction and estimates its parameters automatically. Fiona is able to correct substitution, insertion and deletion errors and can be applied to any sequencing technology. It uses an efficient implementation of the partial suffix array to detect read overlaps with different seed lengths in parallel. We tested Fiona on several real datasets from a variety of organisms with different read lengths and compared its performance with state-of-the-art methods. Fiona shows a constantly higher correction accuracy over a broad range of datasets from 454 and Ion Torrent sequencers, without compromise in speed. CONCLUSION: Fiona is an accurate parameter-free read error-correction method that can be run on inexpensive hardware and can make use of multicore parallelization whenever available. Fiona was implemented using the SeqAn library for sequence analysis and is publicly available for download at http://www.seqan.de/projects/fiona. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marcel H. Schulz, David Weese, Manuel Holtgrewe, Viktoria Dimitrova, Sijia Niu, Knut Reinert, Hugues Richard |
Bioinform. | 3 |
| 2014 | Genome alignment with graph data structures: a comparisonabstractBACKGROUND: Recent advances in rapid, low-cost sequencing have opened up the opportunity to study complete genome sequences. The computational approach of multiple genome alignment allows investigation of evolutionarily related genomes in an integrated fashion, providing a basis for downstream analyses such as rearrangement studies and phylogenetic inference.Graphs have proven to be a powerful tool for coping with the complexity of genome-scale sequence alignments. The potential of graphs to intuitively represent all aspects of genome alignments led to the development of graph-based approaches for genome alignment. These approaches construct a graph from a set of local alignments, and derive a genome alignment through identification and removal of graph substructures that indicate errors in the alignment. RESULTS: We compare the structures of commonly used graphs in terms of their abilities to represent alignment information. We describe how the graphs can be transformed into each other, and identify and classify graph substructures common to one or more graphs. Based on previous approaches, we compile a list of modifications that remove these substructures. CONCLUSION: We show that crucial pieces of alignment information, associated with inversions and duplications, are not visible in the structure of all graphs. If we neglect vertex or edge labels, the graphs differ in their information content. Still, many ideas are shared among all graph-based approaches. Based on these findings, we outline a conceptual framework for graph-based genome alignment that can assist in the development of future genome alignment tools. Birte Kehr, Kathrin Trappe, Manuel Holtgrewe, Knut Reinert |
BMC Bioinform. | 3 |
| 2012 | RazerS 3: Faster, fully sensitive read mappingabstractMOTIVATION: During the past years, next-generation sequencing has become a key technology for many applications in the biomedical sciences. Throughput continues to increase and new protocols provide longer reads than currently available. In almost all applications, read mapping is a first step. Hence, it is crucial to have algorithms and implementations that perform fast, with high sensitivity, and are able to deal with long reads and a large absolute number of insertions and deletions. RESULTS: RazerS is a read mapping program with adjustable sensitivity based on counting q-grams. In this work, we propose the successor RazerS 3, which now supports shared-memory parallelism, an additional seed-based filter with adjustable sensitivity, a much faster, banded version of the Myers' bit-vector algorithm for verification, memory-saving measures and support for the SAM output format. This leads to a much improved performance for mapping reads, in particular, long reads with many errors. We extensively compare RazerS 3 with other popular read mappers and show that its results are often superior to them in terms of sensitivity while exhibiting practical and often competitive run times. In addition, RazerS 3 works without a pre-computed index. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are freely available for download at http://www.seqan.de/projects/razers. RazerS 3 is implemented in C++ and OpenMP under a GPL license using the SeqAn library and supports Linux, Mac OS X and Windows. David Weese, Manuel Holtgrewe, Knut Reinert |
Bioinform. | 2 |
| 2011 | A Novel And Well-Defined Benchmarking Method For Second Generation Read MappingabstractBACKGROUND: Second generation sequencing technologies yield DNA sequence data at ultra high-throughput. Common to most biological applications is a mapping of the reads to an almost identical or highly similar reference genome. The assessment of the quality of read mapping results is not straightforward and has not been formalized so far. Hence, it has not been easy to compare different read mapping approaches in a unified way and to determine which program is the best for what task. RESULTS: We present a new benchmark method, called Rabema (Read Alignment BEnchMArk), for read mappers. It consists of a strict definition of the read mapping problem and of tools to evaluate the result of arbitrary read mappers supporting the SAM output format. CONCLUSIONS: We show the usefulness of the benchmark program by performing a comparison of popular read mappers. The tools supporting the benchmark are licensed under the GPL and available from http://www.seqan.de/projects/rabema.html. Manuel Holtgrewe, Anne-Katrin Emde, David Weese, Knut Reinert |
BMC Bioinform. | 1 |
| 2010 | Simple and Fast Nearest Neighbor SearchabstractWe present a simple randomized data structure for two-dimensional point sets that allows fast nearest neighbor queries in many cases. An implementation outperforms several previous implementations for commonly used benchmarks. Marcel Birn, Manuel Holtgrewe, Peter Sanders 0001, Johannes Singler |
ALENEX | 2 |
| 2010 | Engineering a scalable high quality graph partitionerabstractWe describe an approach to parallel graph partitioning that scales to hundreds of processors and produces a high solution quality. For example, for many instances from Walshaw's benchmark collection we improve the best known partitioning. We use the well known framework of multi-level graph partitioning. All components are implemented by scalable parallel algorithms. Quality improvements compared to previous systems are due to better prioritization of edges to be contracted, better approximation algorithms for identifying matchings, better local search heuristics, and perhaps most notably, a parallelization of the FM local search algorithm that works more locally than previous approaches. Manuel Holtgrewe, Peter Sanders 0001, Christian Schulz 0003 |
IPDPS | 1 |