Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Manuel Holtgrewe

dblp:64/7481 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0002-3051-1763ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection
0.612022
ClearCNV: CNV calling from NGS panel data in the presence of ambiguity and noise · Bioinform. 2022
Bioinformatics and computational biology › genomics
variant calling
0.612022
ClearCNV: CNV calling from NGS panel data in the presence of ambiguity and noise · Bioinform. 2022
Bioinformatics and computational biology › bioinformatics infrastructure
sequencing data management
0.412020
DigestiFlow: from BCL to FASTQ with ease · Bioinform. 2020
Bioinformatics and computational biology › sequence analysis
sequencing data processing
0.412020
DigestiFlow: from BCL to FASTQ with ease · Bioinform. 2020
Bioinformatics and computational biology
genomics
0.312017
HLA-MA: simple yet powerful matching of samples using HLA typing results · Bioinform. 2017
Bioinformatics and computational biology › genomics
high-throughput sequencing
0.312017
HLA-MA: simple yet powerful matching of samples using HLA typing results · Bioinform. 2017
Bioinformatics and computational biology › immunoinformatics
HLA typing
0.312017
HLA-MA: simple yet powerful matching of samples using HLA typing results · Bioinform. 2017
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly
0.212015
Methods for the detection and assembly of novel sequence in high-throughput sequencing data · Bioinform. 2015
Bioinformatics and computational biology › genomics › structural variation
structural variant detection
0.212015
Methods for the detection and assembly of novel sequence in high-throughput sequencing data · Bioinform. 2015
Bioinformatics and computational biology › sequence analysis
sequencing error correction
0.212014
Fiona: a parallel and automatic strategy for read error correction · Bioinform. 2014
Bioinformatics and computational biology › sequence analysis
read mapping
0.112012
RazerS 3: Faster, fully sensitive read mapping · Bioinform. 2012
Parallel and multicore computing › parallel computing
multicore parallelism
0.112014
Fiona: a parallel and automatic strategy for read error correction · Bioinform. 2014
Parallel and multicore computing
parallel computing
0.112014
Fiona: a parallel and automatic strategy for read error correction · Bioinform. 2014
Algorithms and data structures › sequence algorithms › string algorithms
sequence alignment
0.012012
RazerS 3: Faster, fully sensitive read mapping · Bioinform. 2012
Algorithms and data structures › sequence algorithms
string algorithms
0.012012
RazerS 3: Faster, fully sensitive read mapping · Bioinform. 2012

Methods — techniques the papers use, named apart from their topics

homogeneous subset identification · 0.6enrichment kit assignment · 0.6workflow automation · 0.4statistical error model · 0.4partial suffix array · 0.4parameter estimation · 0.4q-gram counting · 0.3paired-end read analysis · 0.2de novo assembly · 0.2seed-based filter · 0.1myers' bitvector algorithm · 0.1myers' bit-vector algorithm · 0.1
YearPublicationVenuePosition
2022 ClearCNV: CNV calling from NGS panel data in the presence of ambiguity and noise
abstract
MOTIVATION: While the identification of small variants in panel sequencing data can be considered a solved problem, the identification of larger, multi-exon copy number variants (CNVs) still poses a considerable challenge. Thus, CNV calling has not been established in all laboratories performing panel sequencing. At the same time, such laboratories have accumulated large datasets and thus have the need to identify CNVs on their data to close the diagnostic gap. RESULTS: In this article, we present our method clearCNV that addresses this need in two ways. First, it helps laboratories to properly assign datasets to enrichment kits. Based on homogeneous subsets of data, clearCNV identifies CNVs affecting the targeted regions. Using real-world datasets and validation, we show that our method is highly competitive with previous methods and preferable in terms of specificity. AVAILABILITY AND IMPLEMENTATION: The software is available for free under a permissible license at https://github.com/bihealth/clear-cnv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vinzenz May, Leonard Koch, Björn Fischer-Zirnsak, Denise Horn, Petra Gehle, Uwe Kornak, Dieter Beule, Manuel Holtgrewe
Bioinform.8
2020 DigestiFlow: from BCL to FASTQ with ease
abstract
Abstract Summary Management of raw-sequencing data and its pre-processing (conversion into sequences and demultiplexing) remains a challenging topic for groups running sequencing devices. They face many challenges in such efforts and solutions ranging from manual management of spreadsheets to very complex and customized laboratory information management systems handling much more than just sequencing raw data. In this article, we describe the software package DigestiFlow that focuses on the management of Illumina flow cell sample sheets and raw data. It allows for automated extraction of information from flow cell data and management of sample sheets. Furthermore, it allows for the automated and reproducible conversion of Illumina base calls to sequences and the demultiplexing thereof using bcl2fastq and Picard Tools, followed by quality control report generation. Availability and implementation The software is available under the MIT license at https://github.com/bihealth/digestiflow-server. The client software components are available via Bioconda. Supplementary information Supplementary data are available at Bioinformatics online.
Manuel Holtgrewe, Clemens Messerschmidt, Mikko Nieminen, Dieter Beule
Bioinform.1
2017 HLA-MA: simple yet powerful matching of samples using HLA typing results
abstract
SUMMARY: We propose the simple method HLA-MA for consistency checking in pipelines operating on human HTS data. The method is based on the HLA typing result of the state-of-the-art method OptiType. Provided that there is sufficient coverage of the HLA loci, comparing HLA types allows for simple, fast and robust matching of samples from whole genome, exome and RNA-seq data. Our approach uses information from small but genetically highly variable regions and thus complements approaches that rely on genome or exon-wide variant profiles. AVAILABILITY AND IMPLEMENTATION: The software is implemented In Python 3 and freely available under the MIT license at https://github.com/bihealth/hlama and via Bioconda. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Clemens Messerschmidt, Manuel Holtgrewe, Dieter Beule
Bioinform.2
2015 Methods for the detection and assembly of novel sequence in high-throughput sequencing data
abstract
MOTIVATION: Large insertions of novel sequence are an important type of structural variants. Previous studies used traditional de novo assemblers for assembling non-mapping high-throughput sequencing (HTS) or capillary reads and then tried to anchor them in the reference using paired read information. RESULTS: We present approaches for detecting insertion breakpoints and targeted assembly of large insertions from HTS paired data: BASIL and ANISE. On near identity repeats that are hard for assemblers, ANISE employs a repeat resolution step. This results in far better reconstructions than obtained by the compared methods. On simulated data, we found our insert assembler to be competitive with the de novo assemblers ABYSS and SGA while yielding already anchored inserted sequence as opposed to unanchored contigs as from ABYSS/SGA. On real-world data, we detected novel sequence in a human individual and thoroughly validated the assembled sequence. ANISE was found to be superior to the competing tool MindTheGap on both simulated and real-world data. AVAILABILITY AND IMPLEMENTATION: ANISE and BASIL are available for download at http://www.seqan.de/projects/herbarium under a permissive open source license.
Manuel Holtgrewe, Léon Kuchenbecker, Knut Reinert
Bioinform.1
2014 Fiona: a parallel and automatic strategy for read error correction
abstract
MOTIVATION: Automatic error correction of high-throughput sequencing data can have a dramatic impact on the amount of usable base pairs and their quality. It has been shown that the performance of tasks such as de novo genome assembly and SNP calling can be dramatically improved after read error correction. While a large number of methods specialized for correcting substitution errors as found in Illumina data exist, few methods for the correction of indel errors, common to technologies like 454 or Ion Torrent, have been proposed. RESULTS: We present Fiona, a new stand-alone read error-correction method. Fiona provides a new statistical approach for sequencing error detection and optimal error correction and estimates its parameters automatically. Fiona is able to correct substitution, insertion and deletion errors and can be applied to any sequencing technology. It uses an efficient implementation of the partial suffix array to detect read overlaps with different seed lengths in parallel. We tested Fiona on several real datasets from a variety of organisms with different read lengths and compared its performance with state-of-the-art methods. Fiona shows a constantly higher correction accuracy over a broad range of datasets from 454 and Ion Torrent sequencers, without compromise in speed. CONCLUSION: Fiona is an accurate parameter-free read error-correction method that can be run on inexpensive hardware and can make use of multicore parallelization whenever available. Fiona was implemented using the SeqAn library for sequence analysis and is publicly available for download at http://www.seqan.de/projects/fiona. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marcel H. Schulz, David Weese, Manuel Holtgrewe, Viktoria Dimitrova, Sijia Niu, Knut Reinert, Hugues Richard
Bioinform.3
2014 Genome alignment with graph data structures: a comparison
abstract
BACKGROUND: Recent advances in rapid, low-cost sequencing have opened up the opportunity to study complete genome sequences. The computational approach of multiple genome alignment allows investigation of evolutionarily related genomes in an integrated fashion, providing a basis for downstream analyses such as rearrangement studies and phylogenetic inference.Graphs have proven to be a powerful tool for coping with the complexity of genome-scale sequence alignments. The potential of graphs to intuitively represent all aspects of genome alignments led to the development of graph-based approaches for genome alignment. These approaches construct a graph from a set of local alignments, and derive a genome alignment through identification and removal of graph substructures that indicate errors in the alignment. RESULTS: We compare the structures of commonly used graphs in terms of their abilities to represent alignment information. We describe how the graphs can be transformed into each other, and identify and classify graph substructures common to one or more graphs. Based on previous approaches, we compile a list of modifications that remove these substructures. CONCLUSION: We show that crucial pieces of alignment information, associated with inversions and duplications, are not visible in the structure of all graphs. If we neglect vertex or edge labels, the graphs differ in their information content. Still, many ideas are shared among all graph-based approaches. Based on these findings, we outline a conceptual framework for graph-based genome alignment that can assist in the development of future genome alignment tools.
Birte Kehr, Kathrin Trappe, Manuel Holtgrewe, Knut Reinert
BMC Bioinform.3
2012 RazerS 3: Faster, fully sensitive read mapping
abstract
MOTIVATION: During the past years, next-generation sequencing has become a key technology for many applications in the biomedical sciences. Throughput continues to increase and new protocols provide longer reads than currently available. In almost all applications, read mapping is a first step. Hence, it is crucial to have algorithms and implementations that perform fast, with high sensitivity, and are able to deal with long reads and a large absolute number of insertions and deletions. RESULTS: RazerS is a read mapping program with adjustable sensitivity based on counting q-grams. In this work, we propose the successor RazerS 3, which now supports shared-memory parallelism, an additional seed-based filter with adjustable sensitivity, a much faster, banded version of the Myers' bit-vector algorithm for verification, memory-saving measures and support for the SAM output format. This leads to a much improved performance for mapping reads, in particular, long reads with many errors. We extensively compare RazerS 3 with other popular read mappers and show that its results are often superior to them in terms of sensitivity while exhibiting practical and often competitive run times. In addition, RazerS 3 works without a pre-computed index. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are freely available for download at http://www.seqan.de/projects/razers. RazerS 3 is implemented in C++ and OpenMP under a GPL license using the SeqAn library and supports Linux, Mac OS X and Windows.
David Weese, Manuel Holtgrewe, Knut Reinert
Bioinform.2
2011 A Novel And Well-Defined Benchmarking Method For Second Generation Read Mapping
abstract
BACKGROUND: Second generation sequencing technologies yield DNA sequence data at ultra high-throughput. Common to most biological applications is a mapping of the reads to an almost identical or highly similar reference genome. The assessment of the quality of read mapping results is not straightforward and has not been formalized so far. Hence, it has not been easy to compare different read mapping approaches in a unified way and to determine which program is the best for what task. RESULTS: We present a new benchmark method, called Rabema (Read Alignment BEnchMArk), for read mappers. It consists of a strict definition of the read mapping problem and of tools to evaluate the result of arbitrary read mappers supporting the SAM output format. CONCLUSIONS: We show the usefulness of the benchmark program by performing a comparison of popular read mappers. The tools supporting the benchmark are licensed under the GPL and available from http://www.seqan.de/projects/rabema.html.
Manuel Holtgrewe, Anne-Katrin Emde, David Weese, Knut Reinert
BMC Bioinform.1
2010 Simple and Fast Nearest Neighbor Search
abstract
We present a simple randomized data structure for two-dimensional point sets that allows fast nearest neighbor queries in many cases. An implementation outperforms several previous implementations for commonly used benchmarks.
Marcel Birn, Manuel Holtgrewe, Peter Sanders 0001, Johannes Singler
ALENEX2
2010 Engineering a scalable high quality graph partitioner
abstract
We describe an approach to parallel graph partitioning that scales to hundreds of processors and produces a high solution quality. For example, for many instances from Walshaw's benchmark collection we improve the best known partitioning. We use the well known framework of multi-level graph partitioning. All components are implemented by scalable parallel algorithms. Quality improvements compared to previous systems are due to better prioritization of edges to be contracted, better approximation algorithms for identifying matchings, better local search heuristics, and perhaps most notably, a parallelization of the FM local search algorithm that works more locally than previous approaches.
Manuel Holtgrewe, Peter Sanders 0001, Christian Schulz 0003
IPDPS1