VLDB 2026 Research / reviewers in the wild / expert
Kay Nieselt
dblp:29/2209 · also Katja Nieselt, Kay Katja Nieselt, Kay Nieselt-Struwe
· DBLP profile ↗
32ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0002-1283-7065ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scanpath Classification with an n-mer Deep Neural Network Architecture
Wolfgang Fuhl, Susanne Zabel, Kay Nieselt |
ETRA | 3 |
| 2025 | Approaching the holistic transcriptome - convolution and deconvolution in transcriptomicsabstractTissues, organs, and entire organisms are composed of diverse cell populations, which are characterized by cell-type-specific gene activities. Bulk RNA-seq represents a robust, cost-effective, scalable method to measure gene activity at the bulk tissue level. However, pathomolecular processes lead to divergent changes in tissue composition and cell-type-specific gene deregulations, which cannot be resolved at the tissue bulk level without information on either change in cell-type proportion or expression at the single-cell level. Accordingly, methods have been developed that constrain bulk deconvolution by information from single-cell expression or cell-type proportion. In parallel, convolution methods have been developed to project single-cell expression to bulk tissue level (pseudobulk simulation). In the present review, we provide an overview of existing convolution and deconvolution methods, their interconnectivity, and benchmarking. Our unique approach lies in the joint consideration of both directions in a "holistic transcriptome model." Through analysis of published (de)convolution studies and benchmarks, we identified the reduced availability of suitable datasets and the use of inaccurate convolution-like methods for (de)convolution model assessment and training as key bottlenecks in the field. On that basis, we conclude with a holistic transcriptome model envisioning that a more integral approach to convolution and deconvolution is needed. With our suggestions for a unified framework we aim to spark collaborative efforts to enable major leaps forward in the field of (de)convolution. Maik Wolfram-Schauerte, Thomas Vogel 0008, Hanati Tuoken, Maria Faelth Savitski, Eric Simon, Kay Nieselt |
Briefings Bioinform. | 6 |
| 2024 | ProtEGOnist: Visual Analysis of Interactions in Small World Networks Using Ego-graphsabstractAbstract Visualizing small‐world networks such as protein‐protein interaction networks or social networks often leads to visual clutter and limited interpretability. To overcome these problems, we presentProtEGOnist, a visualization approach designed to explore small‐world networks.ProtEGOnistvisualizes networks using ego‐graphs that represent local neighborhoods. Ego‐graphs are visualized in an aggregated state as a glyph where the size encodes the size of the neighborhood and in a detailed version where the original network nodes can be explored. The ego‐graphs are arranged in an ego‐graph network, where edges encode similarity using the Jaccard index. Our design aims to reduce visual complexity and clutter while enabling detailed exploration and facilitating the discovery of meaningful patterns. To achieve this, our approach offers a network overview using ego‐graphs, a radar chart for a one‐to‐many ego‐graph comparison and meta‐data integration, and detailed ego‐graph subnetworks for interactive exploration. We demonstrate the applicability of our approach on a co‐author network and two different protein‐protein interaction networks. A web‐based prototype ofProtEGOnistcan be accessed online at https://protegonist-tuevis.cs.uni-tuebingen.de/ . Nicolas Brich, Theresa Anisja Harbig, Mathias Witte Paz, Kay Nieselt, Michael Krone |
Comput. Graph. Forum | 4 |
| 2024 | VIPurPCA: Visualizing and Propagating Uncertainty in Principal Component AnalysisabstractVariables obtained by experimental measurements or statistical inference typically carry uncertainties. When an algorithm uses such quantities as input variables, this uncertainty should propagate to the algorithm's output. Concretely, we consider the classic notion of principal component analysis (PCA): If it is applied to a finite data matrix containing imperfect (i.e., uncertain) multidimensional measurements, its output-a lower-dimensional representation-is itself subject to uncertainty. We demonstrate that this uncertainty can be approximated by appropriate linearization of the algorithm's nonlinear functionality, using automatic differentiation. By itself, however, this structured, uncertain output is difficult to interpret for users. We provide an animation method that effectively visualizes the uncertainty of the lower dimensional map. Implemented as an open-source software package, it allows researchers to assess the reliability of PCA embeddings. Susanne Zabel, Philipp Hennig, Kay Nieselt |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | A temporally quantized distribution of pupil diameters as a new feature for cognitive load classificationabstractIn this paper, we present a new feature that can be used to classify cognitive load based on pupil information. The feature consists of a temporal segmentation of the eye tracking recordings. For each segment of the temporal partition, a probability distribution of pupil size is computed and stored. These probability distributions can then be used to classify the cognitive load. The presented feature significantly improves the classification accuracy of the cognitive load compared to other statistical values obtained from eye tracking data, which represent the state of the art in this field. The applications of determining Cognitive Load from pupil data are numerous and could lead, for example, to pre-warning systems for burnouts. Wolfgang Fuhl, Anne Herrmann-Werner, Kay Nieselt |
ETRA | 3 |
| 2023 | The Tiny Eye Movement TransformerabstractIn this paper, we evaluate different small neural network models for eye movement classification and show our so far developed improved model architecture. For evaluation, we used a subset (1.5 million sequences) of the TEyeDS annotations since it contains in the wild recordings and has the most eye movement annotations to our knowledge. We classified fixations, saccades, and smooth pursuits with four different network architectures and the proposed model improves the equally weighted accuracy by 3.8% to the best competitor while only using 6% of the amount of learnable weights. Wolfgang Fuhl, Anne Herrmann-Werner, Kay Nieselt |
ETRA | 3 |
| 2023 | Area of interest adaption using feature importanceabstractIn this paper, we present two approaches and algorithms that adapt areas of interest (AOI) or regions of interest (ROI), respectively, to the eye tracking data quality and classification task. The first approach uses feature importance in a greedy way and grows or shrinks AOIs in all directions. The second approach is an extension of the first approach, which divides the AOIs into areas and calculates a direction of growth, i.e. a gradient. Both approaches improve the classification results considerably in the case of generalized AOIs, but can also be used for qualitative analysis. In qualitative analysis, the algorithms presented allow the AOIs to be adapted to the data, which means that errors and inaccuracies in eye tracking data can be better compensated for. A good application example is abstract art, where manual AOIs annotation is hardly possible, and data-driven approaches are mainly used for initial AOIs. Wolfgang Fuhl, Susanne Zabel, Theresa Anisja Harbig, Julia Astrid Moldt, Teresa Festl-Wietek, Anne Herrmann-Werner, Kay Nieselt |
ETRA | 7 |
| 2023 | One step closer to EEG based eye trackingabstractIn this paper, we present two approaches and algorithms that adapt areas of interest. We present a new deep neural network (DNN) that can be used to directly determine gaze position using EEG data. EEG-based eye tracking is a new and difficult research topic in the field of eye tracking, but it provides an alternative to image-based eye tracking with an input data set comparable to conventional image processing. The presented DNN exploits spatial dependencies of the EEG signal and uses convolutions similar to spatial filtering, which is used for preprocessing EEG signals. By this, we improve the direct gaze determination from the EEG signal compared to the state of the art by 3.5 cm MAE (Mean absolute error), but unfortunately still do not achieve a directly applicable system, since the inaccuracy is still significantly higher compared to image-based eye trackers. Wolfgang Fuhl, Susanne Zabel, Theresa Anisja Harbig, Julia Astrid Moldt, Teresa Festl-Wietek, Anne Herrmann-Werner, Kay Nieselt |
ETRA | 7 |
| 2023 | GO-Compass: Visual Navigation of Multiple Lists of GO termsabstractAbstract Analysis pipelines in genomics, transcriptomics, and proteomics commonly produce lists of genes, e.g., differentially expressed genes. Often these lists overlap only partly or not at all and contain too many genes for manual comparison. However, using background knowledge, such as the functional annotations of the genes, the lists can be abstracted to functional terms. One approach is to run Gene Ontology (GO) enrichment analyses to determine over‐ and/or underrepresented functions for every list of genes. Due to the hierarchical structure of the Gene Ontology, lists of enriched GO terms can contain many closely related terms, rendering the lists still long, redundant, and difficult to interpret for researchers. In this paper, we present GO‐Compass (Gene Ontology list comparison using Semantic Similarity), a visual analytics tool for the dispensability reduction and visual comparison of lists of GO terms. For dispensability reduction, we adapted the RE‐VIGO algorithm, a summarization method based on the semantic similarity of GO terms, to perform hierarchical dispensability clustering on multiple lists. In an interactive dashboard, GO‐Compass offers several visualizations for the comparison and improved interpretability of GO terms lists. The hierarchical dispensability clustering is visualized as a tree, where users can interactively filter out dispensable GO terms and create flat clusters by cutting the tree at a chosen dispensability. The flat clusters are visualized in animated treemaps and are compared using a correlation heatmap, UpSet plots, and bar charts. With two use cases on published datasets from different omics domains, we demonstrate the general applicability and effectiveness of our approach. In the first use case, we show how the tool can be used to compare lists of differentially expressed genes from a transcriptomics pipeline and incorporate gene information into the analysis. In the second use case using genomics data, we show how GO‐Compass facilitates the analysis of many hundreds of GO terms. For qualitative evaluation of the tool, we conducted feedback sessions with five domain experts and received positive comments. GO‐Compass is part of the Tue‐Vis Visualization Server as a web application available at https://go‐compass‐tuevis.cs.uni‐tuebingen.de/ Theresa Anisja Harbig, Mathias Witte Paz, Kay Nieselt |
Comput. Graph. Forum | 3 |
| 2022 | Foreword
Kay Nieselt, Steffen Oeltze-Jafra, Thomas Schultz 0001, Noeska N. Smit, Björn Sommer 0001 |
Comput. Graph. | 1 |
| 2021 | DamageProfiler: fast damage pattern calculation for ancient DNAabstractMOTIVATION: In ancient DNA research, the authentication of ancient samples based on specific features remains a crucial step in data analysis. Because of this central importance, researchers lacking deeper programming knowledge should be able to run a basic damage authentication analysis. Such software should be user-friendly and easy to integrate into an analysis pipeline. RESULTS: DamageProfiler is a Java-based, stand-alone software to determine damage patterns in ancient DNA. The results are provided in various file formats and plots for further processing. DamageProfiler has an intuitive graphical as well as command line interface that allows the tool to be easily embedded into an analysis pipeline. AVAILABILITY AND IMPLEMENTATION: All of the source code is freely available on GitHub (https://github.com/Integrative-Transcriptomics/DamageProfiler). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Judith Neukamm, Alexander Peltzer, Kay Nieselt |
Bioinform. | 3 |
| 2021 | Foreword: Special section on the Eurographics Workshop on Visual Computing for Biology and Medicine (EG VCBM) 2020
Barbora Kozlíková, Michael Krone, Kay Nieselt, Renata G. Raidou, Noeska N. Smit |
Comput. Graph. | 3 |
| 2019 | Efficient merging of genome profile alignmentsabstractMOTIVATION: Whole-genome alignment (WGA) methods show insufficient scalability toward the generation of large-scale WGAs. Profile alignment-based approaches revolutionized the fields of multiple sequence alignment construction methods by significantly reducing computational complexity and runtime. However, WGAs need to consider genomic rearrangements between genomes, which make the profile-based extension of several whole-genomes challenging. Currently, none of the available methods offer the possibility to align or extend WGA profiles. RESULTS: Here, we present genome profile alignment, an approach that aligns the profiles of WGAs and that is capable of producing large-scale WGAs many times faster than conventional methods. Our concept relies on already available whole-genome aligners, which are used to compute several smaller sets of aligned genomes that are combined to a full WGA with a divide and conquer approach. To align or extend WGA profiles, we make use of the SuperGenome data structure, which features a bidirectional mapping between individual sequence and alignment coordinates. This data structure is used to efficiently transfer different coordinate systems into a common one based on the principles of profiles alignments. The approach allows the computation of a WGA where alignments are subsequently merged along a guide tree. The current implementation uses progressiveMauve and offers the possibility for parallel computation of independent genome alignments. Our results based on various bacterial datasets up to several hundred genomes show that we can reduce the runtime from months to hours with a quality that is negligibly worse than the WGA computed with the conventional progressiveMauve tool. AVAILABILITY AND IMPLEMENTATION: GPA is freely available at https://lambda.informatik.uni-tuebingen.de/gitlab/ahennig/GPA. GPA is implemented in Java, uses progressiveMauve and offers a parallel computation of WGAs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. André Hennig, Kay Nieselt |
Bioinform. | 2 |
| 2015 | Highlights from the 5th Symposium on Biological Data Visualization: Part 1abstractHigh-throughput and high-resolution experimental methods in biology pose enormous challenges for current biological data visualization approaches. To address these challenges, researchers in the visualization and bioinformatics communities need to engage in the design, implementation, application, and evaluation of novel visualization techniques and tools that provide insight into large and highly complex data sets.
BioVis 2015 - the fifth Symposium on Biological Data Visualization - brought together researchers from the visualization, bioinformatics, and biology communities to establish an interdisciplinary dialogue and promote the sharing of expertise between both meeting participants and the communities at large. The meeting educated, inspired, and engaged visualization researchers in problems in biological data visualization as well as bioinformatics and biology researchers in state-of-the-art visualization research. The symposium serves as a platform for researchers from these fields to increase the impact of data visualization approaches in biology. The BioVis 2015 symposium is affiliated with ISMB, the Intelligent Systems for Molecular Biology conference, as a Special Interest Group (SIG) and was colocated with ISMB in Dublin, Ireland, July 10-11 2015.
Each paper was reviewed by researchers from both the bioinformatics and visualization fields and was evaluated for improvements over state-of-the-art and for scientific soundness. The review process was organized in two review cycles. In the first review cycle, each paper was reviewed by three to four reviewers. In the second review cycle, the primary reviewers checked whether the required revisions for conditionally accepted papers were successfully included. Based on the reviewers' scores, reviews, and recommendations, the BioVis 2015 Paper and Publication Chairs and the BMC Bioinformatics Section Editor together selected those that would be published as a BMC Bioinformatics supplement.
The papers from BioVis 2015 appear in two different proceedings: As of the 5th Symposium on Biological Data Visualization: Part 1 in this BMC Bioinformatics supplement and as of the 5th Symposium on Biological Data Visualization: Part 2 in BMC Proceedings (http://www.biomedcentral.com/bmcproc/supplements/9/S6). From the 21 papers submitted to BioVis 2015, 9 papers are published in this BMC Bioinformatics supplement and 5 papers are published in BMC Proceedings.
The articles in this supplement cover a wide spectrum of challenging problems in biological data visualization and their solutions. Overall, three main themes arise from the BioVis 2015 articles: omics, proteins, and imaging. In the omics field, Younesy et al. [1] describe VisRseq: a user-friendly interface for biologists to use libraries in R that provides a method for linking R-apps with interactive components. Chelaru et al. [2] expand on the design behind Epiviz, another tool for bringing genome visualization and computational environments together. Hennig et al. [3] describe Pan-Tetris and Aurisano et al. [4] describe BactoGeNIE: both systems are designed for comparing different genomes. The XCluSim tool by L'Yi et al. [5] has a more general application field and aims to provide insight into how different clustering results relate to each other. In the protein field, Stolte et al. [6] give an overview of the design decisions that underlie Aquaria, a visual analytics tool for exploring protein-related data. Finally, three papers are included from the imaging field. Topics range from image generation, as discussed by Abdellah et al. [7], to a method for parameter optimization in image processing by Pretorius et al. [9] (e.g. for cell nuclei detection and colour deconvolution for histology), and all the way to graph-based exploration of histology images in the GRAPHIE system proposed by Ding et al. [8].
The diversity of topics covered in this issue highlights the wide range of challenges in applying existing visualization techniques to biological data. With this analysis and formalization of our collective experiences, we hope to motivate visualization researchers to think about new problems and new approaches to pressing problems in biology. Jan Aerts, G. Elisabeta Marai, Kay Nieselt, Cydney B. Nielsen, Marc Streit, Daniel Weiskopf |
BMC Bioinform. | 3 |
| 2015 | Pan-Tetris: an interactive visualisation for Pan-genomesabstractBACKGROUND: Large-scale genome projects have paved the way to microbial pan-genome analyses. Pan-genomes describe the union of all genes shared by all members of the species or taxon under investigation. They offer a framework to assess the genomic diversity of a given collection of individual genomes and moreover they help to consolidate gene predictions and annotations. The computation of pan-genomes is often a challenge, and many techniques that use a global alignment-independent approach run the risk of not separating paralogs from orthologs. Also alignment-based approaches which take the gene neighbourhood into account often need additional manual curation of the results. This is quite time consuming and so far there is no visualisation tool available that offers an interactive GUI for the pan-genome to support curating pan-genomic computations or annotations of orthologous genes. RESULTS: We introduce Pan-Tetris, a Java based interactive software tool that provides a clearly structured and suitable way for the visual inspection of gene occurrences in a pan-genome table. The main features of Pan-Tetris are a standard coordinate based presentation of multiple genomes complemented by easy to use tools compensating for algorithmic weaknesses in the pan-genome generation workflow. We demonstrate an application of Pan-Tetris to the pan-genome of Staphylococcus aureus. CONCLUSIONS: Pan-Tetris is currently the only interactive pan-genome visualisation tool. Pan-Tetris is available from http://bit.ly/1vVxYZT. André Hennig, Jörg Bernhardt, Kay Nieselt |
BMC Bioinform. | 3 |
| 2014 | inPHAP: Interactive visualization of genotype and phased haplotype dataabstractBACKGROUND: To understand individual genomes it is necessary to look at the variations that lead to changes in phenotype and possibly to disease. However, genotype information alone is often not sufficient and additional knowledge regarding the phase of the variation is needed to make correct interpretations. Interactive visualizations, that allow the user to explore the data in various ways, can be of great assistance in the process of making well informed decisions. But, currently there is a lack for visualizations that are able to deal with phased haplotype data. RESULTS: We present inPHAP, an interactive visualization tool for genotype and phased haplotype data. inPHAP features a variety of interaction possibilities such as zooming, sorting, filtering and aggregation of rows in order to explore patterns hidden in large genetic data sets. As a proof of concept, we apply inPHAP to the phased haplotype data set of Phase 1 of the 1000 Genomes Project. Thereby, inPHAP's ability to show genetic variations on the population as well as on the individuals level is demonstrated for several disease related loci. CONCLUSIONS: As of today, inPHAP is the only visual analytical tool that allows the user to explore unphased and phased haplotype data interactively. Due to its highly scalable design, inPHAP can be applied to large datasets with up to 100 GB of data, enabling users to visualize even large scale input data. inPHAP closes the gap between common visualization tools for unphased genotype data and introduces several new features, such as the visualization of phased data. inPHAP is available for download at http://bit.ly/1iJgKmX. Günter Jäger, Alexander Peltzer, Kay Nieselt |
BMC Bioinform. | 3 |
| 2013 | Visualizing dimensionality reduction of systems biology data
Andreas M. Lehrmann, Michael Huber 0002, Aydin Can Polatkan, Albert Pritzkau, Kay Nieselt |
Data Min. Knowl. Discov. | 5 |
| 2012 | GenomeRing: alignment visualization based on SuperGenome coordinatesabstractMOTIVATION: The number of completely sequenced genomes is continuously rising, allowing for comparative analyses of genomic variation. Such analyses are often based on whole-genome alignments to elucidate structural differences arising from insertions, deletions or from rearrangement events. Computational tools that can visualize genome alignments in a meaningful manner are needed to help researchers gain new insights into the underlying data. Such visualizations typically are either realized in a linear fashion as in genome browsers or by using a circular approach, where relationships between genomic regions are indicated by arcs. Both methods allow for the integration of additional information such as experimental data or annotations. However, providing a visualization that still allows for a quick and comprehensive interpretation of all important genomic variations together with various supplemental data, which may be highly heterogeneous, remains a challenge. RESULTS: Here, we present two complementary approaches to tackle this problem. First, we propose the SuperGenome concept for the computation of a common coordinate system for all genomes in a multiple alignment. This coordinate system allows for the consistent placement of genome annotations in the presence of insertions, deletions and rearrangements. Second, we present the GenomeRing visualization that, based on the SuperGenome, creates an interactive overview visualization of the multiple genome alignment in a circular layout. We demonstrate our methods by applying them to an alignment of Campylobacter jejuni strains for the discovery of genomic islands as well as to an alignment of Helicobacter pylori, which we visualize in combination with gene expression data. AVAILABILITY: GenomeRing and example data is available at http://it.inf.uni-tuebingen.de/software/genomering/. Alexander Herbig, Günter Jäger, Florian Battke, Kay Nieselt |
Bioinform. | 4 |
| 2012 | Reveal - visual eQTL analyticsabstractMOTIVATION: The analysis of expression quantitative trait locus (eQTL) data is a challenging scientific endeavor, involving the processing of very large, heterogeneous and complex data. Typical eQTL analyses involve three types of data: sequence-based data reflecting the genotypic variations, gene expression data and meta-data describing the phenotype. Based on these, certain genotypes can be connected with specific phenotypic outcomes to infer causal associations of genetic variation, expression and disease. To this end, statistical methods are used to find significant associations between single nucleotide polymorphisms (SNPs) or pairs of SNPs and gene expression. A major challenge lies in summarizing the large amount of data as well as statistical results and to generate informative, interactive visualizations. RESULTS: We present Reveal, our visual analytics approach to this challenge. We introduce a graph-based visualization of associations between SNPs and gene expression and a detailed genotype view relating summarized patient cohort genotypes with data from individual patients and statistical analyses. AVAILABILITY: Reveal is included in Mayday, our framework for visual exploration and analysis. It is available at http://it.inf.uni-tuebingen.de/software/reveal/. CONTACT: [email protected]. Günter Jäger, Florian Battke, Kay Nieselt |
Bioinform. | 3 |
| 2012 | An eQTL biological data visualization challenge and approaches from the visualization communityabstractIn 2011, the IEEE VisWeek conferences inaugurated a symposium on Biological Data Visualization. Like other domain-oriented Vis symposia, this symposium's purpose was to explore the unique characteristics and requirements of visualization within the domain, and to enhance both the Visualization and Bio/Life-Sciences communities by pushing Biological data sets and domain understanding into the Visualization community, and well-informed Visualization solutions back to the Biological community. Amongst several other activities, the BioVis symposium created a data analysis and visualization contest. Unlike many contests in other venues, where the purpose is primarily to allow entrants to demonstrate tour-de-force programming skills on sample problems with known solutions, the BioVis contest was intended to whet the participants' appetites for a tremendously challenging biological domain, and simultaneously produce viable tools for a biological grand challenge domain with no extant solutions. For this purpose expression Quantitative Trait Locus (eQTL) data analysis was selected. In the BioVis 2011 contest, we provided contestants with a synthetic eQTL data set containing real biological variation, as well as a spiked-in gene expression interaction network influenced by single nucleotide polymorphism (SNP) DNA variation and a hypothetical disease model. Contestants were asked to elucidate the pattern of SNPs and interactions that predicted an individual's disease state. 9 teams competed in the contest using a mixture of methods, some analytical and others through visual exploratory methods. Independent panels of visualization and biological experts judged entries. Awards were given for each panel's favorite entry, and an overall best entry agreed upon by both panels. Three special mention awards were given for particularly innovative and useful aspects of those entries. And further recognition was given to entries that correctly answered a bonus question about how a proposed "gene therapy" change to a SNP might change an individual's disease status, which served as a calibration for each approaches' applicability to a typical domain question. In the future, BioVis will continue the data analysis and visualization contest, maintaining the philosophy of providing new challenging questions in open-ended and dramatically underserved Bio/Life Sciences domains. Christopher W. Bartlett, Soo Yeon Cheong, Liping Hou, Jesse Paquette, Pek Yee Lum, Günter Jäger, Florian Battke, Corinna Vehlow, Julian Heinrich, Kay Nieselt, Ryo Sakai, Jan Aerts, William C. Ray |
BMC Bioinform. | 10 |
| 2012 | iHAT: interactive Hierarchical Aggregation Table for Genetic Association DataabstractIn the search for single-nucleotide polymorphisms which influence the observable phenotype, genome wide association studies have become an important technique for the identification of associations between genotype and phenotype of a diverse set of sequence-based data. We present a methodology for the visual assessment of single-nucleotide polymorphisms using interactive hierarchical aggregation techniques combined with methods known from traditional sequence browsers and cluster heatmaps. Our tool, the interactive Hierarchical Aggregation Table (iHAT), facilitates the visualization of multiple sequence alignments, associated metadata, and hierarchical clusterings. Different color maps and aggregation strategies as well as filtering options support the user in finding correlations between sequences and metadata. Similar to other visualizations such as parallel coordinates or heatmaps, iHAT relies on the human pattern-recognition ability for spotting patterns that might indicate correlation or anticorrelation. We demonstrate iHAT using artificial and real-world datasets for DNA and protein association studies as well as expression Quantitative Trait Locus data. Julian Heinrich, Corinna Vehlow, Florian Battke, Günter Jäger, Daniel Weiskopf, Kay Nieselt |
BMC Bioinform. | 6 |
| 2011 | GaggleBridge: collaborative data analysisabstractMOTIVATION: Tools aiding in collaborative data analysis are becoming ever more important as researchers work together over long distances. We present an extension to the Gaggle framework, which has been widely adopted as a tool to enable data exchange between different analysis programs on one computer. RESULTS: Our program, GaggleBridge, transparently extends this functionality to allow data exchange between Gaggle users at different geographic locations using network communication. GaggleBridge can automatically set up SSH tunnels to traverse firewalls while adding some security features to the Gaggle communication. AVAILABILITY: GaggleBridge is available as open-source software implemented in the Java language at http://it.inf.uni-tuebingen.de/gb. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Florian Battke, Stephan Symons, Alexander Herbig, Kay Nieselt |
Bioinform. | 4 |
| 2011 | Identifying associations between amino acid changes and meta information in alignmentsabstractMOTIVATION: We present a method that identifies associations between amino acid changes in potentially significant sites in an alignment (taking into account several amino acid properties) with phenotypic data, through the phylogenetic mixed model. The latter accounts for the dependency of the observations (organisms). It is known from previous studies that the pathogenic aspect of many organisms may be associated with a single or just few changes in amino acids, which have a strong structural and/or functional impact on the protein. Discovering these sites is a big step toward understanding pathogenicity. Our method is able to discover such sites in proteins responsible for the pathogenic character of a group of bacteria. RESULTS: We use our method to predict potentially significant sites in the RpoS protein from a set of 209 bacteria. Several sites with significant differences in biological relevant regions were found. AVAILABILITY: Our tool is publicly available on the CRAN network at http://cran.r-project.org/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lucía Spangenberg, Florian Battke, Martí Graña, Kay Nieselt, Hugo Naya |
Bioinform. | 4 |
| 2011 | MGV: a generic graph viewer for comparative omics dataabstractMOTIVATION: High-throughput transcriptomics, proteomics and metabolomics methods have revolutionized our knowledge of biological systems. To gain knowledge from comparative omics studies, strong data integration and visualization features are required. Knowledge gained from these studies is often available in the form of graphs, and their visualization is especially useful in a wide range of systems biology topics, including pathway analysis, interaction networks or gene models. Especially, it is necessary to compare biological models with measured data. This allows the identification of new models and new insights into existing ones. RESULTS: We present MGV, a versatile generic graph viewer for multiomics data. MGV is integrated into Mayday (Battke et al., 2010). It extends Mayday's visual analytics capabilities by integrating a wide range of biological models, high-throughput data and meta information to display enriched graphs that combine data and models. A wide range of tools is available for visualization of nodes, data-aware graph layout as well as automatic and manual aggregation and refinement of the data. We show the usefulness of MGV applied to several problems, including differential expression of alternative transcripts, transcription factor interaction, cross-study clustering comparison and integration of transcriptomics and metabolomics data for pathway analysis. AVAILABILITY: MGV is a open-source software implemented in Java and freely available as a part of Mayday at www.microarray-analysis.org/mayday. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Stephan Symons, Kay Nieselt |
Bioinform. | 2 |
| 2011 | nocoRNAc: Characterization of non-coding RNAs in prokaryotesabstractBACKGROUND: The interest in non-coding RNAs (ncRNAs) constantly rose during the past few years because of the wide spectrum of biological processes in which they are involved. This led to the discovery of numerous ncRNA genes across many species. However, for most organisms the non-coding transcriptome still remains unexplored to a great extent. Various experimental techniques for the identification of ncRNA transcripts are available, but as these methods are costly and time-consuming, there is a need for computational methods that allow the detection of functional RNAs in complete genomes in order to suggest elements for further experiments. Several programs for the genome-wide prediction of functional RNAs have been developed but most of them predict a genomic locus with no indication whether the element is transcribed or not. RESULTS: We present NOCORNAc, a program for the genome-wide prediction of ncRNA transcripts in bacteria. NOCORNAc incorporates various procedures for the detection of transcriptional features which are then integrated with functional ncRNA loci to determine the transcript coordinates. We applied RNAz and NOCORNAc to the genome of Streptomyces coelicolor and detected more than 800 putative ncRNA transcripts most of them located antisense to protein-coding regions. Using a custom design microarray we profiled the expression of about 400 of these elements and found more than 300 to be transcribed, 38 of them are predicted novel ncRNA genes in intergenic regions. The expression patterns of many ncRNAs are similarly complex as those of the protein-coding genes, in particular many antisense ncRNAs show a high expression correlation with their protein-coding partner. CONCLUSIONS: We have developed NOCORNAc, a framework that facilitates the automated characterization of functional ncRNAs. NOCORNAc increases the confidence of predicted ncRNA loci, especially if they contain transcribed ncRNAs. NOCORNAc is not restricted to intergenic regions, but it is applicable to the prediction of ncRNA transcripts in whole microbial genomes. The software as well as a user guide and example data is available at http://www.zbit.uni-tuebingen.de/pas/nocornac.htm. Alexander Herbig, Kay Nieselt |
BMC Bioinform. | 2 |
| 2010 | Mayday - integrative analytics for expression dataabstractBACKGROUND: DNA Microarrays have become the standard method for large scale analyses of gene expression and epigenomics. The increasing complexity and inherent noisiness of the generated data makes visual data exploration ever more important. Fast deployment of new methods as well as a combination of predefined, easy to apply methods with programmer's access to the data are important requirements for any analysis framework. Mayday is an open source platform with emphasis on visual data exploration and analysis. Many built-in methods for clustering, machine learning and classification are provided for dissecting complex datasets. Plugins can easily be written to extend Mayday's functionality in a large number of ways. As Java program, Mayday is platform-independent and can be used as Java WebStart application without any installation. Mayday can import data from several file formats, database connectivity is included for efficient data organization. Numerous interactive visualization tools, including box plots, profile plots, principal component plots and a heatmap are available, can be enhanced with metadata and exported as publication quality vector files. RESULTS: We have rewritten large parts of Mayday's core to make it more efficient and ready for future developments. Among the large number of new plugins are an automated processing framework, dynamic filtering, new and efficient clustering methods, a machine learning module and database connectivity. Extensive manual data analysis can be done using an inbuilt R terminal and an integrated SQL querying interface. Our visualization framework has become more powerful, new plot types have been added and existing plots improved. CONCLUSIONS: We present a major extension of Mayday, a very versatile open-source framework for efficient micro array data analysis designed for biologists and bioinformaticians. Most everyday tasks are already covered. The large number of available plugins as well as the extension possibilities using compiled plugins and ad-hoc scripting allow for the rapid adaption of Mayday also to very specialized data exploration. Mayday is available at http://microarray-analysis.org. Florian Battke, Stephan Symons, Kay Nieselt |
BMC Bioinform. | 3 |
| 2009 | Prequips - an extensible software platform for integration, visualization and analysis of LC-MS/MS proteomics dataabstractSUMMARY: We describe an integrative software platform, Prequips, for comparative proteomics-based systems biology analysis that: (i) integrates all information generated from mass spectrometry (MS)-based proteomics as well as from basic proteomics data analysis tools, (ii) visualizes such information for various proteomic analyses via graphical interfaces and (iii) links peptide and protein abundances to external tools often used in systems biology studies. AVAILABILITY: http://prequips.sourceforge.net Nils Gehlenborg, Wei Yan 0033, Inyoul Y. Lee, Hyuntae Yoo, Kay Nieselt, Daehee Hwang, Ruedi Aebersold, Leroy Hood |
Bioinform. | 5 |
| 2008 | Post-Hybridization Quality Measures for Oligos in Genome-Wide Microarray Experiments
Florian Battke, Carsten Müller-Tidow, Hubert Serve, Kay Nieselt |
WABI | 4 |
| 2006 | Mayday-a microarray data analysis workbenchabstractUNLABELLED: Mayday is a workbench for visualization, analysis and storage of microarray data. It features a graphical user interface and supports the development and integration of existing and new analysis methods. Besides the infrastructural core functionality, Mayday offers a variety of plug-ins, such as various interactive viewers, a connection to the R statistical environment, a connection to SQL-based databases and different data mining methods, including WEKA-library based methods for classification and various clustering methods. In addition, so-called meta information objects are provided for annotation of the microarray data allowing integration of data from different sources, which is a feature that, for instance, is employed in the enhanced heatmap visualization. SUPPLEMENTARY INFORMATION: The software and more detailed information including screenshots and a user guide as well as test data can be found on the Mayday home page http://www.zbit.uni-tuebingen.de/pas/mayday. The core is published under the GPL (GNU Public License) and the associated plug-ins under the LGPL (Lesser GNU Public License). Janko Dietzsch, Nils Gehlenborg, Kay Nieselt |
Bioinform. | 3 |
| 2005 | Whole-genome prokaryotic phylogenyabstractCurrent understanding of the phylogeny of prokaryotes is based on the comparison of the highly conserved small ssu-rRNA subunit and similar regions. Although such molecules have proved to be very useful phylogenetic markers, mutational saturation is a problem, due to their restricted lengths. Now, a growing number of complete prokaryotic genomes are available. This paper addresses the problem of determining a prokaryotic phylogeny utilizing the comparison of complete genomes. We introduce a new strategy, GBDP, 'genome blast distance phylogeny', and show that different variants of this approach robustly produce phylogenies that are biologically sound, when applied to 91 prokaryotic genomes. In this approach, first Blast is used to compare genomes, then a distance matrix is computed, and finally a tree- or network-reconstruction method such as UPGMA, Neighbor-Joining, BioNJ or Neighbor-Net is applied. Stefan R. Henz, Daniel H. Huson, Alexander F. Auch, Kay Nieselt, Stephan C. Schuster |
Bioinform. | 4 |
| 2004 | DIALIGN P: Fast pair-wise and multiple sequence alignment using parallel processorsabstractBACKGROUND: Parallel computing is frequently used to speed up computationally expensive tasks in Bioinformatics. RESULTS: Herein, a parallel version of the multi-alignment program DIALIGN is introduced. We propose two ways of dividing the program into independent sub-routines that can be run on different processors: (a) pair-wise sequence alignments that are used as a first step to multiple alignment account for most of the CPU time in DIALIGN. Since alignments of different sequence pairs are completely independent of each other, they can be distributed to multiple processors without any effect on the resulting output alignments. (b) For alignments of large genomic sequences, we use a heuristics by splitting up sequences into sub-sequences based on a previously introduced anchored alignment procedure. For our test sequences, this combined approach reduces the program running time of DIALIGN by up to 97%. CONCLUSIONS: By distributing sub-routines to multiple processors, the running time of DIALIGN can be crucially improved. With these improvements, it is possible to apply the program in large-scale genomics and proteomics projects that were previously beyond its scope. Martin Schmollinger, Kay Nieselt, Michael Kaufmann 0001, Burkhard Morgenstern |
BMC Bioinform. | 2 |
| 2003 | Distance Corrections on Recombinant Sequences
David Bryant, Daniel H. Huson, Tobias H. Klöpper, Kay Nieselt |
WABI | 4 |