Daniel H. Huson

dblp:03/1116 · DBLP profile ↗
← Back
53ranked-venue papers
24as first author
7since 2021 · last 2025
0000-0002-2961-604XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 45 · 24 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021Theory of computation · 2Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Sketch, capture and layout phylogenies
abstract
Phylogenetic trees and networks play a central role in biology, bioinformatics, and mathematical biology, and producing clear, informative visualizations of them is an important task. We present new algorithms for visualizing rooted phylogenetic networks in either a "combining" or "transfer" view, in both cladogram and phylogram style. In addition, we introduce a layout algorithm that aims to improve clarity by minimizing the total reticulate displacement of reticulate edges. To address the common issue that biological publications often omit machine-readable representations of depicted trees and networks, we also provide an image-based algorithm that assists in extracting their topology from figures. All algorithms are implemented in our new open source PhyloSketch app.
Daniel H. Huson
PLoS Comput. Biol.1
2024 Leverage the Explainability of Transformer Models to Improve the DNA 5-Methylcytosine Identification (Student Abstract)
abstract
DNA methylation is an epigenetic mechanism for regulating gene expression, and it plays an important role in many biological processes. While methylation sites can be identified using laboratory techniques, much work is being done on developing computational approaches using machine learning. Here, we present a deep-learning algorithm for determining the 5-methylcytosine status of a DNA sequence. We propose an ensemble framework that treats the self-attention score as an explicit feature that is added to the encoder layer generated by fine-tuned language models. We evaluate the performance of the model under different data distribution scenarios.
Wenhuan Zeng, Daniel H. Huson
AAAI2
2024 CatReNet: interactive analysis of (auto-) catalytic reaction networks
abstract
SUMMARY: Catalytic reaction networks serve as fundamental models for understanding biochemical systems. CatReNet is a novel software designed to facilitate interactive analysis of such networks. It offers fast and exact algorithms for computing various types of self-sustaining autocatalytic subnetworks, including so-called CAFs (constructively autocatalytic food-generated networks), RAFs (reflexively autocatalytic food-generated networks), and pseudo-RAFs. It provides dynamic visualizations to aid exploration and understanding. AVAILABILITY AND IMPLEMENTATION: This open-source Java application runs on Linux, MacOS, and Windows. It is available at https://github.com/husonlab/catrenet under a GPL3 license.
Daniel H. Huson, Joana C. Xavier, Mike A. Steel
Bioinform.1
2023 Microbiome Metabolome Integration Platform (MMIP): a web-based platform for microbiome and metabolome data integration and feature identification
abstract
A microbial community maintains its ecological dynamics via metabolite crosstalk. Hence, knowledge of the metabolome, alongside its populace, would help us understand the functionality of a community and also predict how it will change in atypical conditions. Methods that employ low-cost metagenomic sequencing data can predict the metabolic potential of a community, that is, its ability to produce or utilize specific metabolites. These, in turn, can potentially serve as markers of biochemical pathways that are associated with different communities. We developed MMIP (Microbiome Metabolome Integration Platform), a web-based analytical and predictive tool that can be used to compare the taxonomic content, diversity variation and the metabolic potential between two sets of microbial communities from targeted amplicon sequencing data. MMIP is capable of highlighting statistically significant taxonomic, enzymatic and metabolic attributes as well as learning-based features associated with one group in comparison with another. Furthermore, MMIP can predict linkages among species or groups of microbes in the community, specific enzyme profiles, compounds or metabolites associated with such a group of organisms. With MMIP, we aim to provide a user-friendly, online web server for performing key microbiome-associated analyses of targeted amplicon sequencing data, predicting metabolite signature, and using learning-based linkage analysis, without the need for initial metabolomic analysis, and thereby helping in hypothesis generation.
Anupam Gautam, Debaleena Bhowmik, Sayantani Basu, Wenhuan Zeng, Abhishake Lahiri, Daniel H. Huson, Sandip Paul
Briefings Bioinform.6
2023 MeganServer: facilitating interactive access to metagenomic data on a server
abstract
MOTIVATION: Metagenomic projects often involve large numbers of large sequencing datasets (totaling hundreds of gigabytes of data). Thus, computational preprocessing and analysis are usually performed on a server. The results of such analyses are then usually explored interactively. One approach is to use MEGAN, an interactive program that allows analysis and comparison of metagenomic datasets. Previous releases have required that the user first download the computed data from the server, an increasingly time-consuming process. Here, we present MeganServer, a stand-alone program that serves MEGAN files to the web, using a RESTful API, facilitating interactive analysis in MEGAN, without requiring prior download of the data. We describe a number of different application scenarios. AVAILABILITY AND IMPLEMENTATION: MeganServer is provided as a stand-alone program tools/megan-server in the MEGAN software suite, available at https://software-ab.cs.uni-tuebingen.de/download/megan6. Source is available at: https://github.com/husonlab/megan-ce/tree/master/src/megan/ms. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anupam Gautam, Wenhuan Zeng, Daniel H. Huson
Bioinform.3
2022 DeepToA: an ensemble deep-learning approach to predicting the theater of activity of a microbiome
abstract
MOTIVATION: Metagenomics is the study of microbiomes using DNA sequencing. A microbiome consists of an assemblage of microbes that is associated with a 'theater of activity' (ToA). An important question is, to what degree does the taxonomic and functional content of the former depend on the (details of the) latter? Here, we investigate a related technical question: Given a taxonomic and/or functional profile estimated from metagenomic sequencing data, how to predict the associated ToA? We present a deep-learning approach to this question. We use both taxonomic and functional profiles as input. We apply node2vec to embed hierarchical taxonomic profiles into numerical vectors. We then perform dimension reduction using clustering, to address the sparseness of the taxonomic data and thus make the problem more amenable to deep-learning algorithms. Functional features are combined with textual descriptions of protein families or domains. We present an ensemble deep-learning framework DeepToA for predicting the ToA of amicrobial community, based on taxonomic and functional profiles. We use SHAP (SHapley Additive exPlanations) values to determine which taxonomic and functional features are important for the prediction. RESULTS: Based on 7560 metagenomic profiles downloaded from MGnify, classified into 10 different theaters of activity, we demonstrate that DeepToA has an accuracy of 98.30%. We show that adding textual information to functional features increases the accuracy. AVAILABILITY AND IMPLEMENTATION: Our approach is available at http://ab.inf.uni-tuebingen.de/software/deeptoa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenhuan Zeng, Anupam Gautam, Daniel H. Huson
Bioinform.3
2021 Tegula - exploring a galaxy of two-dimensional periodic tilings
abstract
Periodic tilings play a role in the decorative arts, in construction and in crystal structures. Combinatorial tiling theory allows the systematic generation, visualization and exploration of such tilings of the plane, sphere and hyperbolic plane, using advanced algorithms and software. Here we present a “galaxy” of tilings that consists of the set of all 2.4 billion different types of periodic tilings that have Dress complexity up to 24. We make these available in a database and provide a new program called Tegula that can be used to search and visualize such tilings.
Rüdiger Zeller, Olaf Delgado-Friedrichs, Daniel H. Huson
Comput. Aided Geom. Des.3
2020 MAIRA- real-time taxonomic and functional analysis of long reads on a laptop
abstract
BACKGROUND: Advances in mobile sequencing devices and laptop performance make metagenomic sequencing and analysis in the field a technologically feasible prospect. However, metagenomic analysis pipelines are usually designed to run on servers and in the cloud. RESULTS: MAIRA is a new standalone program for interactive taxonomic and functional analysis of long read metagenomic sequencing data on a laptop, without requiring external resources. The program performs fast, online, genus-level analysis, and on-demand, detailed taxonomic and functional analysis. It uses two levels of frame-shift-aware alignment of DNA reads against protein reference sequences, and then performs detailed analysis using a protein synteny graph. CONCLUSIONS: We envision this software being used by researchers in the field, when access to servers or cloud facilities is difficult, or by individuals that do not routinely access such facilities, such as medical researchers, crop scientists, or teachers.
Benjamin Albrecht, Caner Bagci, Daniel H. Huson
BMC Bioinform.3
2018 Autumn Algorithm - Computation of Hybridization Networks for Realistic Phylogenetic Trees
abstract
A minimum hybridization network is a rooted phylogenetic network that displays two given rooted phylogenetic trees using a minimum number of reticulations. Previous mathematical work on their calculation has usually assumed the input trees to be bifurcating, correctly rooted, or that they both contain the same taxa. These assumptions do not hold in biological studies and "realistic" trees have multifurcations, are difficult to root, and rarely contain the same taxa. We present a new algorithm for computing minimum hybridization networks for a given pair of "realistic" rooted phylogenetic trees. We also describe how the algorithm might be used to improve the rooting of the input trees. We introduce the concept of "autumn trees", a nice framework for the formulation of algorithms based on the mathematics of "maximum acyclic agreement forests". While the main computational problem is hard, the run-time depends mainly on how different the given input trees are. In biological studies, where the trees are reasonably similar, our parallel implementation performs well in practice. The algorithm is available in our open source program Dendroscope 3, providing a platform for biologists to explore rooted phylogenetic networks. We demonstrate the utility of the algorithm using several previously studied data sets.
Daniel H. Huson, Simone Linz
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 RiboTagger: fast and unbiased 16S/18S profiling using whole community shotgun metagenomic or metatranscriptome surveys
abstract
BACKGROUND: Taxonomic profiling of microbial communities is often performed using small subunit ribosomal RNA (SSU) amplicon sequencing (16S or 18S), while environmental shotgun sequencing is often focused on functional analysis. Large shotgun datasets contain a significant number of SSU sequences and these can be exploited to perform an unbiased SSU--based taxonomic analysis. RESULTS: Here we present a new program called RiboTagger that identifies and extracts taxonomically informative ribotags located in a specified variable region of the SSU gene in a high-throughput fashion. CONCLUSIONS: RiboTagger permits fast recovery of SSU-RNA sequences from shotgun nucleic acid surveys of complex microbial communities. The program targets all three domains of life, exhibits high sensitivity and specificity and is substantially faster than comparable programs.
Chin Lui Wesley Goi, Daniel H. Huson, Peter F. R. Little, Rohan B. H. Williams
BMC Bioinform.3
2016 MEGAN Community Edition - Interactive Exploration and Analysis of Large-Scale Microbiome Sequencing Data
abstract
There is increasing interest in employing shotgun sequencing, rather than amplicon sequencing, to analyze microbiome samples. Typical projects may involve hundreds of samples and billions of sequencing reads. The comparison of such samples against a protein reference database generates billions of alignments and the analysis of such data is computationally challenging. To address this, we have substantially rewritten and extended our widely-used microbiome analysis tool MEGAN so as to facilitate the interactive analysis of the taxonomic and functional content of very large microbiome datasets. Other new features include a functional classifier called InterPro2GO, gene-centric read assembly, principal coordinate analysis of taxonomy and function, and support for metadata. The new program is called MEGAN Community Edition (CE) and is open source. By integrating MEGAN CE with our high-throughput DNA-to-protein alignment tool DIAMOND and by providing a new program MeganServer that allows access to metagenome analysis files hosted on a server, we provide a straightforward, yet powerful and complete pipeline for the analysis of metagenome shotgun sequences. We illustrate how to perform a full-scale computational analysis of a metagenomic sequencing project, involving 12 samples and 800 million reads, in less than three days on a single server. All source code is available here: https://github.com/danielhuson/megan-ce.
Daniel H. Huson, Sina Beier, Isabell Flade, Anna Górska, Mohamed El-Hadidi 0001, Suparna Mitra, Hans-Joachim Ruscheweyh, Rewati Tappu
PLoS Comput. Biol.1
2014 A poor man's BLASTX - high-throughput metagenomic protein database search using PAUDA
abstract
SUMMARY: In the context of metagenomics, we introduce a new approach to protein database search called PAUDA, which runs ~10,000 times faster than BLASTX, while achieving about one-third of the assignment rate of reads to KEGG orthology groups, and producing gene and taxon abundance profiles that are highly correlated to those obtained with BLASTX. PAUDA requires <80 CPU hours to analyze a dataset of 246 million Illumina DNA reads from permafrost soil for which a previous BLASTX analysis (on a subset of 176 million reads) reportedly required 800,000 CPU hours, leading to the same clustering of samples by functional profiles. AVAILABILITY: PAUDA is freely available from: http://ab.inf.uni-tuebingen.de/software/pauda. Also supplementary method details are available from this website.
Daniel H. Huson
Bioinform.1
2012 Fast computation of minimum hybridization networks
abstract
MOTIVATION: Hybridization events in evolution may lead to incongruent gene trees. One approach to determining possible interspecific hybridization events is to compute a hybridization network that attempts to reconcile incongruent gene trees using a minimum number of hybridization events. RESULTS: We describe how to compute a representative set of minimum hybridization networks for two given bifurcating input trees, using a parallel algorithm and provide a user-friendly implementation. A simulation study suggests that our program performs significantly better than existing software on biologically relevant data. Finally, we demonstrate the application of such methods in the context of the evolution of the Aegilops/Triticum genera. AVAILABILITY AND IMPLEMENTATION: The algorithm is implemented in the program Dendroscope 3, which is freely available from www.dendroscope.org and runs on all three major operating systems.
Benjamin Albrecht, Céline Scornavacca, Alberto Cenci, Daniel H. Huson
Bioinform.4
2011 Tanglegrams for rooted phylogenetic trees and networks
abstract
MOTIVATION: In systematic biology, one is often faced with the task of comparing different phylogenetic trees, in particular in multi-gene analysis or cospeciation studies. One approach is to use a tanglegram in which two rooted phylogenetic trees are drawn opposite each other, using auxiliary lines to connect matching taxa. There is an increasing interest in using rooted phylogenetic networks to represent evolutionary history, so as to explicitly represent reticulate events, such as horizontal gene transfer, hybridization or reassortment. Thus, the question arises how to define and compute a tanglegram for such networks. RESULTS: In this article, we present the first formal definition of a tanglegram for rooted phylogenetic networks and present a heuristic approach for computing one, called the NN-tanglegram method. We compare the performance of our method with existing tree tanglegram algorithms and also show a typical application to real biological datasets. For maximum usability, the algorithm does not require that the trees or networks are bifurcating or bicombining, or that they are on identical taxon sets. AVAILABILITY: The algorithm is implemented in our program Dendroscope 3, which is freely available from www.dendroscope.org. CONTACT: [email protected]; [email protected].
Céline Scornavacca, Franziska Zickmann, Daniel H. Huson
Bioinform.3
2011 Functional analysis of metagenomes and metatranscriptomes using SEED and KEGG
abstract
BACKGROUND: Metagenomics is the study of microbial organisms using sequencing applied directly to environmental samples. Technological advances in next-generation sequencing methods are fueling a rapid increase in the number and scope of metagenome projects. While metagenomics provides information on the gene content, metatranscriptomics aims at understanding gene expression patterns in microbial communities. The initial computational analysis of a metagenome or metatranscriptome addresses three questions: (1) Who is out there? (2) What are they doing? and (3) How do different datasets compare? There is a need for new computational tools to answer these questions. In 2007, the program MEGAN (MEtaGenome ANalyzer) was released, as a standalone interactive tool for analyzing the taxonomic content of a single metagenome dataset. The program has subsequently been extended to support comparative analyses of multiple datasets. RESULTS: The focus of this paper is to report on new features of MEGAN that allow the functional analysis of multiple metagenomes (and metatranscriptomes) based on the SEED hierarchy and KEGG pathways. We have compared our results with the MG-RAST service for different datasets. CONCLUSIONS: The MEGAN program now allows the interactive analysis and comparison of the taxonomical and functional content of multiple datasets. As a stand-alone tool, MEGAN provides an alternative to web portals for scientists that have concerns about uploading their unpublished data to a website.
Suparna Mitra, Paul Rupek, Daniel C. Richter, Tim Urich, Jack A. Gilbert, Folker Meyer, Andreas Wilke, Daniel H. Huson
BMC Bioinform.8
2010 Phylogenetic networks do not need to be complex: using fewer reticulations to represent conflicting clusters
abstract
UNLABELLED: Phylogenetic trees are widely used to display estimates of how groups of species are evolved. Each phylogenetic tree can be seen as a collection of clusters, subgroups of the species that evolved from a common ancestor. When phylogenetic trees are obtained for several datasets (e.g. for different genes), then their clusters are often contradicting. Consequently, the set of all clusters of such a dataset cannot be combined into a single phylogenetic tree. Phylogenetic networks are a generalization of phylogenetic trees that can be used to display more complex evolutionary histories, including reticulate events, such as hybridizations, recombinations and horizontal gene transfers. Here, we present the new Cass algorithm that can combine any set of clusters into a phylogenetic network. We show that the networks constructed by Cass are usually simpler than networks constructed by other available methods. Moreover, we show that Cass is guaranteed to produce a network with at most two reticulations per biconnected component, whenever such a network exists. We have implemented Cass and integrated it into the freely available Dendroscope software. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Leo van Iersel, Steven Kelk, Regula Rupp, Daniel H. Huson
Bioinform.4
2010 Short clones or long clones? A simulation study on the use of paired reads in metagenomics
abstract
BACKGROUND: Metagenomics is the study of environmental samples using sequencing. Rapid advances in sequencing technology are fueling a vast increase in the number and scope of metagenomics projects. Most metagenome sequencing projects so far have been based on Sanger or Roche-454 sequencing, as only these technologies provide long enough reads, while Illumina sequencing has not been considered suitable for metagenomic studies due to a short read length of only 35 bp. However, now that reads of length 75 bp can be sequenced in pairs, Illumina sequencing has become a viable option for metagenome studies. RESULTS: This paper addresses the problem of taxonomical analysis of paired reads. We describe a new feature of our metagenome analysis software MEGAN that allows one to process sequencing reads in pairs and makes assignments of such reads based on the combined bit scores of their matches to reference sequences. Using this new software in a simulation study, we investigate the use of Illumina paired-sequencing in taxonomical analysis and compare the performance of single reads, short clones and long clones. In addition, we also compare against simulated Roche-454 sequencing runs. CONCLUSION: This work shows that paired reads perform better than single reads, as expected, but also, perhaps slightly less obviously, that long clones allow more specific assignments than short ones. A new version of the program MEGAN that explicitly takes paired reads into account is available from our website.
Suparna Mitra, Max Schubach, Daniel H. Huson
BMC Bioinform.3
2010 New common ancestor problems in trees and directed acyclic graphs
Johannes Fischer 0001, Daniel H. Huson
Inf. Process. Lett.2
2009 Computing galled networks from real data
abstract
MOTIVATION: Developing methods for computing phylogenetic networks from biological data is an important problem posed by molecular evolution and much work is currently being undertaken in this area. Although promising approaches exist, there are no tools available that biologists could easily and routinely use to compute rooted phylogenetic networks on real datasets containing tens or hundreds of taxa. Biologists are interested in clades, i.e. groups of monophyletic taxa, and these are usually represented by clusters in a rooted phylogenetic tree. The problem of computing an optimal rooted phylogenetic network from a set of clusters, is hard, in general. Indeed, even the problem of just determining whether a given network contains a given cluster is hard. Hence, some researchers have focused on topologically restricted classes of networks, such as galled trees and level-k networks, that are more tractable, but have the practical draw-back that a given set of clusters will usually not possess such a representation. RESULTS: In this article, we argue that galled networks (a generalization of galled trees) provide a good trade-off between level of generality and tractability. Any set of clusters can be represented by some galled network and the question whether a cluster is contained in such a network is easy to solve. Although the computation of an optimal galled network involves successively solving instances of two different NP-complete problems, in practice our algorithm solves this problem exactly on large datasets containing hundreds of taxa and many reticulations in seconds, as illustrated by a dataset containing 279 prokaryotes. AVAILABILITY: We provide a fast, robust and easy-to-use implementation of this work in version 2.0 of our tree-handling software Dendroscope, freely available from http://www.dendroscope.org.
Daniel H. Huson, Regula Rupp, Vincent Berry, Philippe Gambette, Christophe Paul
Bioinform.1
2009 Visual and statistical comparison of metagenomes
abstract
BACKGROUND: Metagenomics is the study of the genomic content of an environmental sample of microbes. Advances in the through-put and cost-efficiency of sequencing technology is fueling a rapid increase in the number and size of metagenomic datasets being generated. Bioinformatics is faced with the problem of how to handle and analyze these datasets in an efficient and useful way. One goal of these metagenomic studies is to get a basic understanding of the microbial world both surrounding us and within us. One major challenge is how to compare multiple datasets. Furthermore, there is a need for bioinformatics tools that can process many large datasets and are easy to use. RESULTS: This article describes two new and helpful techniques for comparing multiple metagenomic datasets. The first is a visualization technique for multiple datasets and the second is a new statistical method for highlighting the differences in a pairwise comparison. We have developed implementations of both methods that are suitable for very large datasets and provide these in Version 3 of our standalone metagenome analysis tool MEGAN. CONCLUSION: These new methods are suitable for the visual comparison of many large metagenomes and the statistical comparison of two metagenomes at a time. Nevertheless, more work needs to be done to support the comparative analysis of multiple metagenome datasets. AVAILABILITY: Version 3 of MEGAN, which implements all ideas presented in this article, can be obtained from our web site at: www-ab.informatik.uni-tuebingen.de/software/megan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Suparna Mitra, Bernhard Klar, Daniel H. Huson
Bioinform.3
2009 Methods for comparative metagenomics
abstract
BACKGROUND: Metagenomics is a rapidly growing field of research that aims at studying uncultured organisms to understand the true diversity of microbes, their functions, cooperation and evolution, in environments such as soil, water, ancient remains of animals, or the digestive system of animals and humans. The recent development of ultra-high throughput sequencing technologies, which do not require cloning or PCR amplification, and can produce huge numbers of DNA reads at an affordable cost, has boosted the number and scope of metagenomic sequencing projects. Increasingly, there is a need for new ways of comparing multiple metagenomics datasets, and for fast and user-friendly implementations of such approaches. RESULTS: This paper introduces a number of new methods for interactively exploring, analyzing and comparing multiple metagenomic datasets, which will be made freely available in a new, comparative version 2.0 of the stand-alone metagenome analysis tool MEGAN. CONCLUSION: There is a great need for powerful and user-friendly tools for comparative analysis of metagenomic data and MEGAN 2.0 will help to fill this gap.
Daniel H. Huson, Daniel C. Richter, Suparna Mitra, Alexander F. Auch, Stephan C. Schuster
BMC Bioinform.1
2009 Drawing Rooted Phylogenetic Networks
abstract
The evolutionary history of a collection of species is usually represented by a phylogenetic tree. Sometimes, phylogenetic networks are used as a means of representing reticulate evolution or of showing uncertainty and incompatibilities in evolutionary datasets. This is often done using unrooted phylogenetic networks such as split networks, due in part, to the availability of software (SplitsTree) for their computation and visualization. In this paper we discuss the problem of drawing rooted phylogenetic networks as cladograms or phylograms in a number of different views that are commonly used for rooted trees. Implementations of the algorithms are available in new releases of the Dendroscope and SplitsTree programs.
Daniel H. Huson
IEEE ACM Trans. Comput. Biol. Bioinform.1
2009 Special Section: Phylogenetics
abstract
The seven papers in this special section focus on phylogenetics and these four themes: new data types and algorithms in phylogenetics; reticulate evolution; constructing large trees; and mathematical modeling of evolution.
Daniel H. Huson, Vincent Moulton, Mike A. Steel
IEEE ACM Trans. Comput. Biol. Bioinform.1
2008 Summarizing Multiple Gene Trees Using Cluster Networks
Daniel H. Huson, Regula Rupp
WABI1
2008 Improved Layout of Phylogenetic Networks
abstract
Split networks are increasingly being used in phylogenetic analysis. Usually, a simple equal angle algorithm is used to draw such networks, producing layouts that leave much room for improvement. Addressing the problem of producing better layouts of split networks, this paper presents an algorithm for maximizing the area covered by the network, describes an extension of the equal-daylight algorithm to networks, looks into using a spring embedder and discusses how to construct rooted split networks.
Philippe Gambette, Daniel H. Huson
IEEE ACM Trans. Comput. Biol. Bioinform.2
2007 etagenome Analysis using Megan
Daniel H. Huson, Alexander F. Auch, Stephan C. Schuster
APBC1
2007 Beyond Galled Trees - Decomposition and Computation of Galled Networks
Daniel H. Huson, Tobias H. Klöpper
RECOMB1
2007 COPYCAT : cophylogenetic analysis tool
abstract
UNLABELLED: We have developed the software CopyCat which provides an easy and fast access to cophylogenetic analyses. It incorporates a wrapper for the program ParaFit, which conducts a statistical test for the presence of congruence between host and parasite phylogenies. CopyCat offers various features, such as the creation of customized host-parasite association data and the computation of phylogenetic host/parasite trees based on the NCBI taxonomy. AVAILABILITY: CopyCat and its manual are freely available at http://www-ab.informatik.uni-tuebingen.de/software/copycat. SUPPLEMENTARY INFORMATION: Results of the real-world example can be found at http://www-ab.informatik.uni-tuebingen.de/software/copycat or Bioinformatics online.
Jan P. Meier-Kolthoff, Alexander F. Auch, Daniel H. Huson, Markus Göker
Bioinform.3
2007 OSLay: optimal syntenic layout of unfinished assemblies
abstract
UNLABELLED: The whole genome shotgun approach to genome sequencing results in a collection of contigs that must be ordered and oriented to facilitate efficient gap closure. We present a new tool OSLay that uses synteny between matching sequences in a target assembly and a reference assembly to layout the contigs (or scaffolds) in the target assembly. The underlying algorithm is based on maximum weight matching. The tool provides an interactive visualization of the computed layout and the result can be imported into the assembly editing tool Consed to support the design of primer pairs for gap closure. MOTIVATION: To enhance efficiency in the gap closure phase of a genome project it is crucial to know which contigs are adjacent in the target genome. Related genome sequences can be used to layout contigs in an assembly. AVAILABILITY: OSLay is freely available from: http://www-ab.informatik.unituebingen.de/software/oslay.
Daniel C. Richter, Stephan C. Schuster, Daniel H. Huson
Bioinform.3
2007 Dendroscope: An interactive viewer for large phylogenetic trees
abstract
BACKGROUND: Research in evolution requires software for visualizing and editing phylogenetic trees, for increasingly very large datasets, such as arise in expression analysis or metagenomics, for example. It would be desirable to have a program that provides these services in an efficient and user-friendly way, and that can be easily installed and run on all major operating systems. Although a large number of tree visualization tools are freely available, some as a part of more comprehensive analysis packages, all have drawbacks in one or more domains. They either lack some of the standard tree visualization techniques or basic graphics and editing features, or they are restricted to small trees containing only tens of thousands of taxa. Moreover, many programs are difficult to install or are not available for all common operating systems. RESULTS: We have developed a new program, Dendroscope, for the interactive visualization and navigation of phylogenetic trees. The program provides all standard tree visualizations and is optimized to run interactively on trees containing hundreds of thousands of taxa. The program provides tree editing and graphics export capabilities. To support the inspection of large trees, Dendroscope offers a magnification tool. The software is written in Java 1.4 and installers are provided for Linux/Unix, MacOS X and Windows XP. CONCLUSION: Dendroscope is a user-friendly program for visualizing and navigating phylogenetic trees, for both small and large datasets.
Daniel H. Huson, Daniel C. Richter, Christian Rausch, Tobias Dezulian, Markus Franz, Regula Rupp
BMC Bioinform.1
2006 Reducing Distortion in Phylogenetic Networks
Daniel H. Huson, Mike A. Steel, Jim Whitfield
WABI1
2006 Identification of plant microRNA homologs
abstract
Abstract Summary: MicroRNAs (miRNAs) are a recently discovered class of non-coding RNAs that regulate gene and protein expression in plants and animals. MiRNAs have so far been identified mostly by specific cloning of small RNA molecules, complemented by computational methods. We present a computational identification approach that is able to identify candidate miRNA homologs in any set of sequences, given a query miRNA. The approach is based on a sequence similarity search step followed by a set of structural filters. Availability: microHARVESTER is offered as a web-service and additionally as source code upon request at Contact: [email protected]
Tobias Dezulian, Michael Remmert, Javier F. Palatnik, Detlef Weigel, Daniel H. Huson
Bioinform.5
2005 Reconstruction of Reticulate Networks from Gene Trees
Daniel H. Huson, Tobias H. Klöpper, Peter J. Lockhart, Mike A. Steel
RECOMB1
2005 Whole-genome prokaryotic phylogeny
abstract
Current understanding of the phylogeny of prokaryotes is based on the comparison of the highly conserved small ssu-rRNA subunit and similar regions. Although such molecules have proved to be very useful phylogenetic markers, mutational saturation is a problem, due to their restricted lengths. Now, a growing number of complete prokaryotic genomes are available. This paper addresses the problem of determining a prokaryotic phylogeny utilizing the comparison of complete genomes. We introduce a new strategy, GBDP, 'genome blast distance phylogeny', and show that different variants of this approach robustly produce phylogenies that are biologically sound, when applied to 91 prokaryotic genomes. In this approach, first Blast is used to compare genomes, then a distance matrix is computed, and finally a tree- or network-reconstruction method such as UPGMA, Neighbor-Joining, BioNJ or Neighbor-Net is applied.
Stefan R. Henz, Daniel H. Huson, Alexander F. Auch, Kay Nieselt, Stephan C. Schuster
Bioinform.2
2004 Phylogenetic Super-networks from Partial Trees
Daniel H. Huson, Tobias Dezulian, Tobias H. Klöpper, Mike A. Steel
WABI1
2004 VisRD--visual recombination detection
abstract
SUMMARY: VisRD, a program for visual recombination detection in a sequence alignment is presented. VisRD is written in Java and is designed to complement the multi-purpose phylogenetic software package SplitsTree4. AVAILABILITY: The software is freely available from http://www.lcb.uu.se/~vmoulton/software/visrd/
Sofia K. Forslund, Daniel H. Huson, Vincent Moulton
Bioinform.2
2004 Phylogenetic trees based on gene content
abstract
UNLABELLED: Comparing gene content between species can be a useful approach for reconstructing phylogenetic trees. In this paper, we derive a maximum-likelihood estimation of evolutionary distance between species under a simple model of gene genesis and gene loss. Using simulated data on a biological tree with 107 taxa (and on a number of randomly generated trees), we compare the accuracy of tree reconstruction using this ML distance measure to an earlier ad hoc distance. We then compare these distance-based approaches to a character-based tree reconstruction method (Dollo parsimony) which seems well suited to the analysis of gene content data. To simplify simulations, we give a formal proof of the well-known 'fact' that the Dollo parsimony score is independent of the choice of root. Our results show a consistent trend, with the character-based method and ML distance measure outperforming the earlier ad hoc distance method. AVAILABILITY: http://www.ab.informatik.uni-tuebingen.de/software/genecontent/welcome_en.html
Daniel H. Huson, Mike A. Steel
Bioinform.1
2004 Constructing Splits Graphs
abstract
Phylogenetic trees correspond one-to-one to compatible systems of splits and so splits play an important role in theoretical and computational aspects of phylogeny. Whereas any tree reconstruction method can be thought of as producing a compatible system of splits, an increasing number of phylogenetic algorithms are available that compute split systems that are not necessarily compatible and, thus, cannot always be represented by a tree. Such methods include the split decomposition, Neighbor-Net, consensus networks, and the Z-closure method. A more general split system of this kind can be represented graphically by a so-called splits graph, which generalizes the concept of a phylogenetic tree. This paper addresses the problem of computing a splits graph for a given set of splits. We have implemented all presented algorithms in a new program called SplitsTree4.
Andreas Dress, Daniel H. Huson
IEEE ACM Trans. Comput. Biol. Bioinform.2
2004 Phylogenetic Super-Networks from Partial Trees
abstract
In practice, one is often faced with incomplete phylogenetic data, such as a collection of partial trees or partial splits. This paper poses the problem of inferring a phylogenetic super-network from such data and provides an efficient algorithm for doing so, called the Z-closure method. Additionally, the questions of assigning lengths to the edges of the network and how to restrict the "dimensionality" of the network are addressed. Applications to a set of five published partial gene trees relating different fungal species and to six published partial gene trees relating different grasses illustrate the usefulness of the method and an experimental study confirms its potential. The method is implemented as a plug-in for the program SplitsTree4.
Daniel H. Huson, Tobias Dezulian, Tobias H. Klöpper, Mike A. Steel
IEEE ACM Trans. Comput. Biol. Bioinform.1
2003 Distance Corrections on Recombinant Sequences
David Bryant, Daniel H. Huson, Tobias H. Klöpper, Kay Nieselt
WABI2
2002 Segment Match Refinement and Applications
Aaron L. Halpern, Daniel H. Huson, Knut Reinert
WABI2
2002 The greedy path-merging algorithm for contig scaffolding
abstract
Given a collection of contigs and mate-pairs. The Contig Scaffolding Problem is to order and orientate the given contigs in a manner that is consistent with as many mate-pairs as possible. This paper describes an efficient heuristic called the greedy-path merging algorithm for solving this problem. The method was originally developed as a key component of the compartmentalized assembly strategy developed at Celera Genomics. This interim approach was used at an early stage of the sequencing of the human genome to produce a preliminary assembly based on preliminary whole genome shotgun data produced at Celera and preliminary human contigs produced by the Human Genome Project.
Daniel H. Huson, Knut Reinert, Eugene W. Myers
J. ACM1
2001 The greedy path-merging algorithm for sequence assembly
abstract
Two different approaches to determining the human genome are currently being pursued: one is the “clone-by-clone” approach, employed by the publicly-funded. Human Genome Project, and the other is the “whole genome shotgun” approach, favored by researchers at Celera Genomics. An interim strategy employed at Celera, called hierarchical assembly, makes use of preliminary data produced by both approaches. This paper introduces the Bactig Ordering Problem, which is a key problem that arises in this context, and presents an efficient heuristic called the greedy path-merginq algorithm that performs well on real data.
Daniel H. Huson, Knut Reinert, Eugene W. Myers
RECOMB1
2001 Comparing Assemblies Using Fragments and Mate-Pairs
Daniel H. Huson, Aaron L. Halpern, Zhongwu Lai, Eugene W. Myers, Knut Reinert, Granger G. Sutton
WABI1
2000 The Conserved Exon Method for Gene Finding
Vineet Bafna, Daniel H. Huson
ISMB2
2000 4-Regular Vertex-Transitive Tilings of E3
Olaf Delgado-Friedrichs, Daniel H. Huson
Discret. Comput. Geom.2
1999 Solving Large Scale Phylogenetic Problems using DCM2
Daniel H. Huson, Lisa Vawter, Tandy J. Warnow
ISMB1
1999 Obtaining highly accurate topology estimates of evolutionary trees from very short sequences
abstract
The evolutionary history of a set of species is represented by a phylogenetic tree, in other words, by a rooted, leaf-labelled tree, where internal nodes represent ancestral species and the leaves represent modern day species. Accurate (or even boundedly inaccurate) topology reconstructions of large and divergent trees has long been considered one of the major challenges in systematic biology. None of the polynomial time methods developed by the theoretical computer science community has been shown to outperform the popular Neighbor-Joining method used by systematic biologists, with respect to topology estimation. (However, preliminary experiments indicate that two new variants of Neighbor-Joining, Bio-NJ and Weighbor, do exhibit improved performance.) In this paper, we present a simple polynomial time method, the Disk-Covering Method (DCM), which boosts the performance of base phylogenetic methods. We analyze the performance of DCM-boosted distance methods under the general Markov mo...
Daniel H. Huson, Scott Nettles, Tandy J. Warnow
RECOMB1
1999 Tiling Space by Platonic Solids, I
Olaf Delgado-Friedrichs, Daniel H. Huson
Discret. Comput. Geom.2
1998 SplitsTree: analyzing and visualizing evolutionary data
abstract
MOTIVATION: Real evolutionary data often contain a number of different and sometimes conflicting phylogenetic signals, and thus do not always clearly support a unique tree. To address this problem, Bandelt and Dress (Adv. Math., 92, 47-05, 1992) developed the method of split decomposition. For ideal data, this method gives rise to a tree, whereas less ideal data are represented by a tree-like network that may indicate evidence for different and conflicting phylogenies. RESULTS: SplitsTree is an interactive program, for analyzing and visualizing evolutionary data, that implements this approach. It also supports a number of distances transformations, the computation of parsimony splits, spectral analysis and bootstrapping.
Daniel H. Huson
Bioinform.1
1998 Two Finiteness Theorems for Periodic Tilings of d-Dimensional Euclidean Space
Nikolai P. Dolbilin, Andreas Dress, Daniel H. Huson
Discret. Comput. Geom.3
1996 Analyzing and Visualizing Sequence and Distance Data Using SplitsTree
Andreas Dress, Daniel H. Huson, Vincent Moulton
Discret. Appl. Math.2
1992 The Classification of Quasi-Regular Polyhedra of Genus 2
Reinhard Franz, Daniel H. Huson
Discret. Comput. Geom.2