Michael Hiller

dblp:87/3358 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
0since 2021 · last 2017
0000-0003-3024-1449ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6Graphics, computer vision, multimedia, augmented reality and games · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
comparative genomics
0.632017
chainCleaner improves genome alignment specificity and sensitivity · Bioinform. 2017
CESAR 2.0 substantially improves speed and accuracy of comparative gene annotation · Bioinform. 2017
Computational discovery of human coding and non-coding transcripts with conserved splice sites · Bioinform. 2011
Bioinformatics and computational biology › sequence alignment
genome alignment
0.422017
chainCleaner improves genome alignment specificity and sensitivity · Bioinform. 2017
CESAR 2.0 substantially improves speed and accuracy of comparative gene annotation · Bioinform. 2017
Bioinformatics and computational biology › genome annotation
gene annotation
0.312017
CESAR 2.0 substantially improves speed and accuracy of comparative gene annotation · Bioinform. 2017
Bioinformatics and computational biology › comparative genomics
genome rearrangement
0.312017
chainCleaner improves genome alignment specificity and sensitivity · Bioinform. 2017
Bioinformatics and computational biology › genome annotation › gene prediction
homology-based gene prediction
0.312017
CESAR 2.0 substantially improves speed and accuracy of comparative gene annotation · Bioinform. 2017

Methods — techniques the papers use, named apart from their topics

whole-genome alignment · 0.3local alignment filtering · 0.3chain-based alignment · 0.3tree-based log-odds scoring · 0.1phylogenetic substitution modeling · 0.1machine learning · 0.1
YearPublicationVenuePosition
2017 Iterative error correction of long sequencing reads maximizes accuracy and improves contig assembly
abstract
Next-generation sequencers such as Illumina can now produce reads up to 300 bp with high throughput, which is attractive for genome assembly. A first step in genome assembly is to computationally correct sequencing errors. However, correcting all errors in these longer reads is challenging. Here, we show that reads with remaining errors after correction often overlap repeats, where short erroneous k-mers occur in other copies of the repeat. We developed an iterative error correction pipeline that runs the previously published String Graph Assembler (SGA) in multiple rounds of k-mer-based correction with an increasing k-mer size, followed by a final round of overlap-based correction. By combining the advantages of small and large k-mers, this approach corrects more errors in repeats and minimizes the total amount of erroneous reads. We show that higher read accuracy increases contig lengths two to three times. We provide SGA-Iteratively Correcting Errors (https://github.com/hillerlab/IterativeErrorCorrection/) that implements iterative error correction by using modules from SGA.
Katrin Sameith, Juliana G. Roscito, Michael Hiller
Briefings Bioinform.3
2017 CESAR 2.0 substantially improves speed and accuracy of comparative gene annotation
abstract
MOTIVATION: Homology-based gene prediction is a powerful concept to annotate newly sequenced genomes. We have previously demonstrated that whole genome alignments can be utilized for accurate comparative coding gene annotation. RESULTS: Here we present CESAR 2.0 that utilizes genome alignments to transfer coding gene annotations from one reference to many other aligned genomes. We show that CESAR 2.0 is 77 times faster and requires 31 times less memory compared to its predecessor. CESAR 2.0 substantially improves the ability to align splice sites that have shifted over larger distances, allowing for precise identification of the exon boundaries in the aligned genome. Finally, CESAR 2.0 supports entire genes, which enables the annotation of joined exons that arose by complete intron deletions. CESAR 2.0 can readily be applied to new genome alignments to annotate coding genes in many other genomes at improved accuracy and without necessitating large-computational resources. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at https://github.com/hillerlab/CESAR2.0. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Virag Sharma, Peter Schwede, Michael Hiller
Bioinform.3
2017 chainCleaner improves genome alignment specificity and sensitivity
abstract
MOTIVATION: Accurate alignments between entire genomes are crucial for comparative genomics. However, computing sensitive and accurate genome alignments is a challenging problem, complicated by genomic rearrangements. RESULTS: Here we present a fast approach, called chainCleaner, that improves the specificity in genome alignments by accurately detecting and removing local alignments that obscure the evolutionary history of genomic rearrangements. Systematic tests on alignments between the human and other vertebrate genomes show that chainCleaner (i) improves the alignment of numerous orthologous genes, (ii) exposes alignments between exons of orthologous genes that were masked before by alignments to pseudogenes, and (iii) recovers hundreds of kilobases in local alignments that otherwise would fall below a minimum score threshold. Our approach has broad applicability to improve the sensitivity and specificity of genome alignments. AVAILABILITY AND IMPLEMENTATION: http://bds.mpi-cbg.de/hillerlab/chainCleaner/ or https://github.com/ucscGenomeBrowser/kent. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hernando G. Suarez Duran, Björn E. Langer, Pradnya Ladde, Michael Hiller
Bioinform.4
2011 Computational discovery of human coding and non-coding transcripts with conserved splice sites
abstract
MOTIVATION: Long non-coding RNAs (lncRNAs) resemble protein-coding mRNAs but do not encode proteins. Most lncRNAs are under lower sequence constraints than protein-coding genes and lack conserved secondary structures, making it hard to predict them computationally. RESULTS: We introduce an approach to predict spliced lncRNAs in vertebrate genomes combining comparative genomics and machine learning. It is based on detecting signatures of characteristic splice site evolution in vertebrate whole genome alignments. First, we predict individual splice sites, then assemble compatible sites into exon candidates, and finally predict multi-exon transcripts. Using a novel method to evaluate typical splice site substitution patterns that explicitly takes the species phylogeny into account, we show that individual splice sites can be accurately predicted. Since our approach relies only on predicted splice sites, it can uncover both coding and non-coding exons. We show that our predicted exons and partial transcripts are mostly non-coding and lack conserved secondary structures. These exons are of particular interest, since existing computational approaches cannot detect them. Transcriptome sequencing data indicate tissue-specific expression patterns of predicted exons and there is evidence that increasing sequencing depth and breadth will validate additional predictions. We also found a significant enrichment of predicted exons that form multi-exon transcript parts, and we experimentally validate such a novel multi-exon gene. Overall, we obtain 336 novel multi-exon transcript predictions from human intergenic regions. Our results indicate the existence of novel human transcripts that are conserved in evolution and our approach contributes to the completion of the human transcript catalog. AVAILABILITY AND IMPLEMENTATION: Predicted human splice sites, exons and gene structures together with a Perl implementation of the tree-based log-odds scoring and a supplementary PDF file containing additional figures and tables are available at: http://www.bioinf.uni-leipzig.de/publications/supplements/10-010. The five experimentally confirmed partial transcript isoforms have been deposited in GenBank under accession numbers HM587422-HM587426.
Dominic Rose, Michael Hiller, Katharina Schutt, Jörg Hackermüller, Rolf Backofen, Peter F. Stadler
Bioinform.2
2010 TassDB2 - A comprehensive database of subtle alternative splicing events
abstract
BACKGROUND: Subtle alternative splicing events involving tandem splice sites separated by a short (2-12 nucleotides) distance are frequent and evolutionarily widespread in eukaryotes, and a major contributor to the complexity of transcriptomes and proteomes. However, these events have been either omitted altogether in databases on alternative splicing, or only the cases of experimentally confirmed alternative splicing have been reported. Thus, a database which covers all confirmed cases of subtle alternative splicing as well as the numerous putative tandem splice sites (which might be confirmed once more transcript data becomes available), and allows to search for tandem splice sites with specific features and download the results, is a valuable resource for targeted experimental studies and large-scale bioinformatics analyses of tandem splice sites. Towards this goal we recently set up TassDB (Tandem Splice Site DataBase, version 1), which stores data about alternative splicing events at tandem splice sites separated by 3 nt in eight species. DESCRIPTION: We have substantially revised and extended TassDB. The currently available version 2 contains extensive information about tandem splice sites separated by 2-12 nt for the human and mouse transcriptomes including data on the conservation of the tandem motifs in five vertebrates. TassDB2 offers a user-friendly interface to search for specific genes or for genes containing tandem splice sites with specific features as well as the possibility to download result datasets. For example, users can search for cases of alternative splicing where the proportion of EST/mRNA evidence supporting the minor isoform exceeds a specific threshold, or where the difference in splice site scores is specified by the user. The predicted impact of each event on the protein is also reported, along with information about being a putative target for the nonsense-mediated decay (NMD) pathway. Links are provided to the UCSC genome browser and other external resources. CONCLUSION: TassDB2, available via http://www.tassdb.info, provides comprehensive resources for researchers interested in both targeted experimental studies and large-scale bioinformatics analyses of short distance tandem splice sites.
Rileen Sinha, Thorsten Lenser, Niels Jahn, Ulrike Gausmann, Swetlana Friedel, Karol Szafranski, Klaus Huse, Philip Rosenstiel, Jochen Hampe, Stefan Schuster, Michael Hiller, Rolf Backofen, Matthias Platzer
BMC Bioinform.11
2008 Improved identification of conserved cassette exons using Bayesian networks
abstract
BACKGROUND: Alternative splicing is a major contributor to the diversity of eukaryotic transcriptomes and proteomes. Currently, large scale detection of alternative splicing using expressed sequence tags (ESTs) or microarrays does not capture all alternative splicing events. Moreover, for many species genomic data is being produced at a far greater rate than corresponding transcript data, hence in silico methods of predicting alternative splicing have to be improved. RESULTS: Here, we show that the use of Bayesian networks (BNs) allows accurate prediction of evolutionary conserved exon skipping events. At a stringent false positive rate of 0.5%, our BN achieves an improved true positive rate of 61%, compared to a previously reported 50% on the same dataset using support vector machines (SVMs). Incorporating several novel discriminative features such as intronic splicing regulatory elements leads to the improvement. Features related to mRNA secondary structure increase the prediction performance, corroborating previous findings that secondary structures are important for exon recognition. Random labelling tests rule out overfitting. Cross-validation on another dataset confirms the increased performance. When using the same dataset and the same set of features, the BN matches the performance of an SVM in earlier literature. Remarkably, we could show that about half of the exons which are labelled constitutive but receive a high probability of being alternative by the BN, are in fact alternative exons according to the latest EST data. Finally, we predict exon skipping without using conservation-based features, and achieve a true positive rate of 29% at a false positive rate of 0.5%. CONCLUSION: BNs can be used to achieve accurate identification of alternative exons and provide clues about possible dependencies between relevant features. The near-identical performance of the BN and SVM when using the same features shows that good classification depends more on features than on the choice of classifier. Conservation based features continue to be the most informative, and hence distinguishing alternative exons from constitutive ones without using conservation based features remains a challenging problem.
Rileen Sinha, Michael Hiller, Rainer Pudimat, Ulrike Gausmann, Matthias Platzer, Rolf Backofen
BMC Bioinform.2
2001 Distortion Minimization with Fast Local Search for Fractal Image Compression
Raouf Hamzaoui, Dietmar Saupe, Michael Hiller
J. Vis. Commun. Image Represent.3
2000 Fast Code Enhancement with Local Search for Fractal Image Compression
abstract
Optimal fractal coding consists of finding, in a finite set of contractive affine mappings, one whose unique fixed point is closest to the original image. Optimal fractal coding is an NP-hard combinatorial optimization problem. Conventional coding is based on a greedy suboptimal algorithm known as collage coding. In a previous study, we proposed a local search algorithm that significantly improves on collage coding. However the algorithm, which requires the computation of many fixed points, is computationally expensive. In this paper we provide techniques that drastically reduce the time complexity of the algorithm.
Raouf Hamzaoui, Dietmar Saupe, Michael Hiller
ICIP3