Alexander Goesmann

dblp:08/1793 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-7086-2568ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 PARANOiD: Pipeline for Automated Read ANalysis of iCLIP Data
abstract
MOTIVATION: RNA-protein interactions play essential roles in every living organism, with RNA transcription, processing, and translation being just a few examples. Therefore, determining the set of RNAs that are bound by individual RNA-binding proteins, as well as the precise location of the interaction, is crucial for biological understanding. CLIP (UV-cross-linking and immunoprecipitation) is a method developed to study these interactions. Several variations of the CLIP protocol have been developed, e.g. iCLIP (individual-nucleotide resolution CLIP), which offers nucleotide-precise resolution of the cross-linking event. RESULTS: PARANOiD is a versatile software for fully automated analysis of iCLIP and iCLIP2 data. It contains all steps necessary for preprocessing, the determination of cross-link locations, and several additional steps, which can be used to detect specific characteristics, e.g. definite distances between cross-link events or identify binding motifs. Additionally, results are visualized as statistical plots for a quick overview and as standardized bioinformatics file formats, which can be used for further analysis steps. AVAILABILITY AND IMPLEMENTATION: PARANOiD is published under the MIT license and is available from https://github.com/patrick-barth/PARANOiD. The documentation is available at https://paranoid.readthedocs.io/en/latest/index.html.
Patrick Barth, Frank Förster, Sebastian Jaenicke, Fabienne Thelen, Oliver Rossbach, Friedemann Weber, Lyudmila Shalamova, Alexander Goesmann
Bioinform.8
2026 Towards FAIR and federated data ecosystems for interdisciplinary research
abstract
Scientific data management is at a critical juncture, driven by exponential data growth, increasing cross-domain dependencies, and a severe reproducibility crisis in modern research. Traditional centralized data management approaches are not only struggling with data volume but also fail to address the fragmentation of research results across domains. This hinders scientific reproducibility and cross-domain collaboration and increases concerns about data sovereignty and governance. This article proposes FAIR and federated Data Ecosystems as an improved architectural pattern for future research data ecosystems. It tries to incorporate the latest advancements in decentralized, distributed systems into existing research infrastructure to promote cross-domain collaboration. Based on established patterns from Data Commons, Data Meshes, and Data Spaces, our approach focuses on a layered architecture that consists of governance, data, service, and application layers. With this, it could be possible to preserve domain-specific expertise and control while facilitating data integration through standardized interfaces and semantic enrichment. Key requirements include adaptive metadata management, simplified user interaction, robust security, and transparent data transactions. Our architecture supports compute-to-data as well as data-to-compute paradigms, implementing a decentralized peer-to-peer network that scales horizontally. This article aims to provide both an impulse for the technical architecture as well as concepts for a governance framework so that FAIR and federated Data Ecosystems could enable researchers to build on existing work while maintaining control over their data and computing resources. This could provide a practical path towards an integrated research infrastructure that respects domain autonomy as well as interoperability requirements.
Sebastian Beyvers, Jannis Schlegel, Lukas Brehm, Maria Hansen, Alexander Goesmann, Frank Förster
PLoS Comput. Biol.5
2025 LegionProfiler: a computational tool for the identification of virulence factors and classification of Legionella pneumophila serogroup 1 isolates
abstract
SUMMARY: Legionella pneumophila has significantly contributed to multiple cases of pneumonia with a high rate of mortality globally. Its ability to exploit host mechanisms through several expressed virulence factors poses challenges for diagnosis, treatment, and outbreak control. To address this, we developed LegionProfiler, a computational tool that swiftly identifies virulence factor protein domains within genome assemblies of Legionella pneumophila serogroup 1 isolates and classifies them into high- or low-virulence groups. LegionProfiler automates the probing of genome assemblies for virulence-associated protein domains and determines the isolate's potential to cause severe pneumonia infection. The LegionProfiler workflow is made available through a user-friendly interface to enhance technical control of infectious sources and adds important insights to the general epidemiology of clinical isolates. It could also support the development of targeted therapeutic strategies that will improve patient treatment. AVAILABILITY AND IMPLEMENTATION: LegionProfiler is freely accessible as a web service at https://legionprofiler.uni-muenster.de, and can also be run locally in a Docker container. The source code can be found at https://imigitlab.uni-muenster.de/heiderlab/legionprofiler or at Zenodo (DOI:10.5281/zenodo.15592325). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Oluwafemi A. Sarumi, Wilhelm Bertrams, Oliver Schwengers, Jan-Paul Herrmann, Torsten Hain, Laurine Kieper, Markus Petzold, Alexander Goesmann, Bernd Schmeck, Dominik Heider
Bioinform.8
2025 <tt>Crypt4GH-JS</tt>: securely storing sensitive data online with client-side encryption
abstract
MOTIVATION AND RESULTS: Crypt4GH-JS is a browser-ready implementation of the Crypt4GH file encryption standard written in JavaScript. While having minimal to no impact on data upload and download throughput this library enables on-the-fly encryption of arbitrary data in web applications, regardless of whether on the client or server side. As development moves more and more toward cloud-native applications, this library represents a significant step forward for flexible data security in the context of opaque cloud storage systems. AVAILABILITY AND IMPLEMENTATION: Crypt4GH-JS can be installed via Node Package Manager (https://www.npmjs.com/package/crypt4gh_js) or through its public GitHub Repository (https://github.com/fathelen/crypt4ghJS), where the source code is available. Crypt4GH-JS can be tested in the browser using our demonstration website, which can be found at: https://fathelen.github.io/crypt4ghJS/.
Fabienne Thelen, Jannis Hochmuth, Sven Griep, Benedikt Schwab, Alexander Goesmann, Frank Förster
Bioinform.5
2024 Curare and GenExVis: a versatile toolkit for analyzing and visualizing RNA-Seq data
abstract
Even though high-throughput transcriptome sequencing is routinely performed in many laboratories, computational analysis of such data remains a cumbersome process often executed manually, hence error-prone and lacking reproducibility. For corresponding data processing, we introduce Curare, an easy-to-use yet versatile workflow builder for analyzing high-throughput RNA-Seq data focusing on differential gene expression experiments. Data analysis with Curare is customizable and subdivided into preprocessing, quality control, mapping, and downstream analysis stages, providing multiple options for each step while ensuring the reproducibility of the workflow. For a fast and straightforward exploration and visualization of differential gene expression results, we provide the gene expression visualizer software GenExVis. GenExVis can create various charts and tables from simple gene expression tables and DESeq2 results without the requirement to upload data or install software packages. In combination, Curare and GenExVis provide a comprehensive software environment that supports the entire data analysis process, from the initial handling of raw RNA-Seq data to the final DGE analyses and result visualizations, thereby significantly easing data processing and subsequent interpretation.
Patrick Blumenkamp, Max Pfister, Sonja Diedrich, Karina Brinkrolf, Sebastian Jaenicke, Alexander Goesmann
BMC Bioinform.6
2022 Prediction of antimicrobial resistance based on whole-genome sequencing and machine learning
abstract
MOTIVATION: Antimicrobial resistance (AMR) is one of the biggest global problems threatening human and animal health. Rapid and accurate AMR diagnostic methods are thus very urgently needed. However, traditional antimicrobial susceptibility testing (AST) is time-consuming, low throughput and viable only for cultivable bacteria. Machine learning methods may pave the way for automated AMR prediction based on genomic data of the bacteria. However, comparing different machine learning methods for the prediction of AMR based on different encodings and whole-genome sequencing data without previously known knowledge remains to be done. RESULTS: In this study, we evaluated logistic regression (LR), support vector machine (SVM), random forest (RF) and convolutional neural network (CNN) for the prediction of AMR for the antibiotics ciprofloxacin, cefotaxime, ceftazidime and gentamicin. We could demonstrate that these models can effectively predict AMR with label encoding, one-hot encoding and frequency matrix chaos game representation (FCGR encoding) on whole-genome sequencing data. We trained these models on a large AMR dataset and evaluated them on an independent public dataset. Generally, RFs and CNNs perform better than LR and SVM with AUCs up to 0.96. Furthermore, we were able to identify mutations that are associated with AMR for each antibiotic. AVAILABILITY AND IMPLEMENTATION: Source code in data preparation and model training are provided at GitHub website (https://github.com/YunxiaoRen/ML-iAMR). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yunxiao Ren, Trinad Chakraborty, Swapnil Doijad, Linda Falgenhauer, Jane Falgenhauer, Alexander Goesmann, Anne-Christin Hauschild, Oliver Schwengers, Dominik Heider
Bioinform.6
2020 ASA3P: An automatic and scalable pipeline for the assembly, annotation and higher-level analysis of closely related bacterial isolates
abstract
Whole genome sequencing of bacteria has become daily routine in many fields. Advances in DNA sequencing technologies and continuously dropping costs have resulted in a tremendous increase in the amounts of available sequence data. However, comprehensive in-depth analysis of the resulting data remains an arduous and time-consuming task. In order to keep pace with these promising but challenging developments and to transform raw data into valuable information, standardized analyses and scalable software tools are needed. Here, we introduce ASA3P, a fully automatic, locally executable and scalable assembly, annotation and analysis pipeline for bacterial genomes. The pipeline automatically executes necessary data processing steps, i.e. quality clipping and assembly of raw sequencing reads, scaffolding of contigs and annotation of the resulting genome sequences. Furthermore, ASA3P conducts comprehensive genome characterizations and analyses, e.g. taxonomic classification, detection of antibiotic resistance genes and identification of virulence factors. All results are presented via an HTML5 user interface providing aggregated information, interactive visualizations and access to intermediate results in standard bioinformatics file formats. We distribute ASA3P in two versions: a locally executable Docker container for small-to-medium-scale projects and an OpenStack based cloud computing version able to automatically create and manage self-scaling compute clusters. Thus, automatic and standardized analysis of hundreds of bacterial genomes becomes feasible within hours. The software and further information is available at: asap.computational.bio.
Oliver Schwengers, Andreas Hoek, Moritz Fritzenwanker, Linda Falgenhauer, Torsten Hain, Trinad Chakraborty, Alexander Goesmann
PLoS Comput. Biol.7
2016 ReadXplorer 2 - detailed read mapping analysis and visualization from one single source
abstract
MOTIVATION: The vast amount of already available and currently generated read mapping data requires comprehensive visualization, and should benefit from bioinformatics tools offering a wide spectrum of analysis functionality from just one source. Appropriate handling of multiple mapped reads during mapping analyses remains an issue that demands improvement. RESULTS: The capabilities of the read mapping analysis and visualization tool ReadXplorer were vastly enhanced. Here, we present an even finer granulated read mapping classification, improving the level of detail for analyses and visualizations. The spectrum of automatic analysis functions has been broadened to include genome rearrangement detection as well as correlation analysis between two mapping data sets. Existing functions were refined and enhanced, namely the computation of differentially expressed genes, the read count and normalization analysis and the transcription start site detection. Additionally, ReadXplorer 2 features a highly improved support for large eukaryotic data sets and a command line version, enabling its integration into workflows. Finally, the new version is now able to display any kind of tabular results from other bioinformatics tools. AVAILABILITY AND IMPLEMENTATION: http://www.readxplorer.org CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Rolf Hilker, Kai Bernd Stadermann, Oliver Schwengers, Evgeny Anisiforov, Sebastian Jaenicke, Bernd Weisshaar, Tobias Zimmermann, Alexander Goesmann
Bioinform.8
2014 ReadXplorer - visualization and analysis of mapped sequences
abstract
MOTIVATION: Fast algorithms and well-arranged visualizations are required for the comprehensive analysis of the ever-growing size of genomic and transcriptomic next-generation sequencing data. RESULTS: ReadXplorer is a software offering straightforward visualization and extensive analysis functions for genomic and transcriptomic DNA sequences mapped on a reference. A unique specialty of ReadXplorer is the quality classification of the read mappings. It is incorporated in all analysis functions and displayed in ReadXplorer's various synchronized data viewers for (i) the reference sequence, its base coverage as (ii) normalizable plot and (iii) histogram, (iv) read alignments and (v) read pairs. ReadXplorer's analysis capability covers RNA secondary structure prediction, single nucleotide polymorphism and deletion-insertion polymorphism detection, genomic feature and general coverage analysis. Especially for RNA-Seq data, it offers differential gene expression analysis, transcription start site and operon detection as well as RPKM value and read count calculations. Furthermore, ReadXplorer can combine or superimpose coverage of different datasets. AVAILABILITY AND IMPLEMENTATION: ReadXplorer is available as open-source software at http://www.readxplorer.org along with a detailed manual.
Rolf Hilker, Kai Bernd Stadermann, Daniel Doppmeier, Jörn Kalinowski, Jens Stoye, Jasmin Straube, Jörn Winnebald, Alexander Goesmann
Bioinform.8
2014 AKE - the Accelerated k-mer Exploration web-tool for rapid taxonomic classification and visualization
abstract
BACKGROUND: With the advent of low cost, fast sequencing technologies metagenomic analyses are made possible. The large data volumes gathered by these techniques and the unpredictable diversity captured in them are still, however, a challenge for computational biology. RESULTS: In this paper we address the problem of rapid taxonomic assignment with small and adaptive data models (< 5 MB) and present the accelerated k-mer explorer (AKE). Acceleration in AKE's taxonomic assignments is achieved by a special machine learning architecture, which is well suited to model data collections that are intrinsically hierarchical. We report classification accuracy reasonably well for ranks down to order, observed on a study on real world data (Acid Mine Drainage, Cow Rumen). CONCLUSION: We show that the execution time of this approach is orders of magnitude shorter than competitive approaches and that accuracy is comparable. The tool is presented to the public as a web application (url: https://ani.cebitec.uni-bielefeld.de/ake/ , username: bmc, password: bmcbioinfo).
Daniel Langenkämper, Alexander Goesmann, Tim W. Nattkemper
BMC Bioinform.2
2013 MeltDB 2.0-advances of the metabolomics software system
abstract
MOTIVATION: The research area metabolomics achieved tremendous popularity and development in the last couple of years. Owing to its unique interdisciplinarity, it requires to combine knowledge from various scientific disciplines. Advances in the high-throughput technology and the consequently growing quality and quantity of data put new demands on applied analytical and computational methods. Exploration of finally generated and analyzed datasets furthermore relies on powerful tools for data mining and visualization. RESULTS: To cover and keep up with these requirements, we have created MeltDB 2.0, a next-generation web application addressing storage, sharing, standardization, integration and analysis of metabolomics experiments. New features improve both efficiency and effectivity of the entire processing pipeline of chromatographic raw data from pre-processing to the derivation of new biological knowledge. First, the generation of high-quality metabolic datasets has been vastly simplified. Second, the new statistics tool box allows to investigate these datasets according to a wide spectrum of scientific and explorative questions. AVAILABILITY: The system is publicly available at https://meltdb.cebitec.uni-bielefeld.de. A login is required but freely available.
Nikolas Kessler, Heiko Neuweger, Anja Bonte, Georg Langenkämper, Karsten Niehaus, Tim W. Nattkemper, Alexander Goesmann
Bioinform.7
2011 Exact and complete short-read alignment to microbial genomes using Graphics Processing Unit programming
abstract
MOTIVATION: The introduction of next-generation sequencing techniques and especially the high-throughput systems Solexa (Illumina Inc.) and SOLiD (ABI) made the mapping of short reads to reference sequences a standard application in modern bioinformatics. Short-read alignment is needed for reference based re-sequencing of complete genomes as well as for gene expression analysis based on transcriptome sequencing. Several approaches were developed during the last years allowing for a fast alignment of short sequences to a given template. Methods available to date use heuristic techniques to gain a speedup of the alignments, thereby missing possible alignment positions. Furthermore, most approaches return only one best hit for every query sequence, thus losing the potentially valuable information of alternative alignment positions with identical scores. RESULTS: We developed SARUMAN (Semiglobal Alignment of short Reads Using CUDA and NeedleMAN-Wunsch), a mapping approach that returns all possible alignment positions of a read in a reference sequence under a given error threshold, together with one optimal alignment for each of these positions. Alignments are computed in parallel on graphics hardware, facilitating an considerable speedup of this normally time-consuming step. Combining our filter algorithm with CUDA-accelerated alignments, we were able to align reads to microbial genomes in time comparable or even faster than all published approaches, while still providing an exact, complete and optimal result. At the same time, SARUMAN runs on every standard Linux PC with a CUDA-compatible graphics accelerator. AVAILABILITY: http://www.cebitec.uni-bielefeld.de/brf/saruman/saruman.html.
Jochen Blom, Tobias Jakobi, Daniel Doppmeier, Sebastian Jaenicke, Jörn Kalinowski, Jens Stoye, Alexander Goesmann
Bioinform.7
2011 Conveyor: a workflow engine for bioinformatic analyses
abstract
MOTIVATION: The rapidly increasing amounts of data available from new high-throughput methods have made data processing without automated pipelines infeasible. As was pointed out in several publications, integration of data and analytic resources into workflow systems provides a solution to this problem, simplifying the task of data analysis. Various applications for defining and running workflows in the field of bioinformatics have been proposed and published, e.g. Galaxy, Mobyle, Taverna, Pegasus or Kepler. One of the main aims of such workflow systems is to enable scientists to focus on analysing their datasets instead of taking care for data management, job management or monitoring the execution of computational tasks. The currently available workflow systems achieve this goal, but fundamentally differ in their way of executing workflows. RESULTS: We have developed the Conveyor software library, a multitiered generic workflow engine for composition, execution and monitoring of complex workflows. It features an open, extensible system architecture and concurrent program execution to exploit resources available on modern multicore CPU hardware. It offers the ability to build complex workflows with branches, loops and other control structures. Two example use cases illustrate the application of the versatile Conveyor engine to common bioinformatics problems. AVAILABILITY: The Conveyor application including client and server are available at http://conveyor.cebitec.uni-bielefeld.de.
Burkhard Linke, Robert Giegerich, Alexander Goesmann
Bioinform.3
2009 Qupe - a Rich Internet Application to take a step forward in the analysis of mass spectrometry-based quantitative proteomics experiments
abstract
MOTIVATION: The goal of present -omics sciences is to understand biological systems as a whole in terms of interactions of the individual cellular components. One of the main building blocks in this field of study is proteomics where tandem mass spectrometry (LC-MS/MS) in combination with isotopic labelling techniques provides a common way to obtain a direct insight into regulation at the protein level. Methods to identify and quantify the peptides contained in a sample are well established, and their output usually results in lists of identified proteins and calculated relative abundance values. The next step is to move ahead from these abstract lists and apply statistical inference methods to compare measurements, to identify genes that are significantly up- or down-regulated, or to detect clusters of proteins with similar expression profiles. RESULTS: We introduce the Rich Internet Application (RIA) Qupe providing comprehensive data management and analysis functions for LC-MS/MS experiments. Starting with the import of mass spectra data the system guides the experimenter through the process of protein identification by database search, the calculation of protein abundance ratios, and in particular, the statistical evaluation of the quantification results including multivariate analysis methods such as analysis of variance or hierarchical cluster analysis. While a data model to store these results has been developed, a well-defined programming interface facilitates the integration of novel approaches. A compute cluster is utilized to distribute computationally intensive calculations, and a web service allows to interchange information with other -omics software applications. To demonstrate that Qupe represents a step forward in quantitative proteomics analysis an application study on Corynebacterium glutamicum has been carried out. AVAILABILITY AND IMPLEMENTATION: Qupe is implemented in Java utilizing Hibernate, Echo2, R and the Spring framework. We encourage the usage of the RIA in the sense of the 'software as a service' concept, maintained on our servers and accessible at the following location: http://qupe.cebitec.uni-bielefeld.de. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Stefan P. Albaum, Heiko Neuweger, Benjamin Fränzel, Sita Lange, Dominik Mertens, Christian Trötschel, Dirk Wolters, Jörn Kalinowski, Tim W. Nattkemper, Alexander Goesmann
Bioinform.10
2009 EDGAR: A software framework for the comparative analysis of prokaryotic genomes
abstract
BACKGROUND: The introduction of next generation sequencing approaches has caused a rapid increase in the number of completely sequenced genomes. As one result of this development, it is now feasible to analyze large groups of related genomes in a comparative approach. A main task in comparative genomics is the identification of orthologous genes in different genomes and the classification of genes as core genes or singletons. RESULTS: To support these studies EDGAR - "Efficient Database framework for comparative Genome Analyses using BLAST score Ratios" - was developed. EDGAR is designed to automatically perform genome comparisons in a high throughput approach. Comparative analyses for 582 genomes across 75 genus groups taken from the NCBI genomes database were conducted with the software and the results were integrated into an underlying database. To demonstrate a specific application case, we analyzed ten genomes of the bacterial genus Xanthomonas, for which phylogenetic studies were awkward due to divergent taxonomic systems. The resultant phylogeny EDGAR provided was consistent with outcomes from traditional approaches performed recently and moreover, it was possible to root each strain with unprecedented accuracy. CONCLUSION: EDGAR provides novel analysis features and significantly simplifies the comparative analysis of related genomes. The software supports a quick survey of evolutionary relationships and simplifies the process of obtaining new biological insights into the differential gene content of kindred genomes. Visualization features, like synteny plots or Venn diagrams, are offered to the scientific community through a web-based and therefore platform independent user interface http://edgar.cebitec.uni-bielefeld.de, where the precomputed data sets can be browsed.
Jochen Blom, Stefan P. Albaum, Daniel Doppmeier, Alfred Pühler, Frank-Jörg Vorhölter, Martha Zakrzewski, Alexander Goesmann
BMC Bioinform.7
2009 TACOA - Taxonomic classification of environmental genomic fragments using a kernelized nearest neighbor approach
abstract
BACKGROUND: Metagenomics, or the sequencing and analysis of collective genomes (metagenomes) of microorganisms isolated from an environment, promises direct access to the "unculturable majority". This emerging field offers the potential to lay solid basis on our understanding of the entire living world. However, the taxonomic classification is an essential task in the analysis of metagenomics data sets that it is still far from being solved. We present a novel strategy to predict the taxonomic origin of environmental genomic fragments. The proposed classifier combines the idea of the k-nearest neighbor with strategies from kernel-based learning. RESULTS: Our novel strategy was extensively evaluated using the leave-one-out cross validation strategy on fragments of variable length (800 bp - 50 Kbp) from 373 completely sequenced genomes. TACOA is able to classify genomic fragments of length 800 bp and 1 Kbp with high accuracy until rank class. For longer fragments > or = 3 Kbp accurate predictions are made at even deeper taxonomic ranks (order and genus). Remarkably, TACOA also produces reliable results when the taxonomic origin of a fragment is not represented in the reference set, thus classifying such fragments to its known broader taxonomic class or simply as "unknown". We compared the classification accuracy of TACOA with the latest intrinsic classifier PhyloPythia using 63 recently published complete genomes. For fragments of length 800 bp and 1 Kbp the overall accuracy of TACOA is higher than that obtained by PhyloPythia at all taxonomic ranks. For all fragment lengths, both methods achieved comparable high specificity results up to rank class and low false negative rates are also obtained. CONCLUSION: An accurate multi-class taxonomic classifier was developed for environmental genomic fragments. TACOA can predict with high reliability the taxonomic origin of genomic fragments as short as 800 bp. The proposed method is transparent, fast, accurate and the reference set can be easily updated as newly sequenced genomes become available. Moreover, the method demonstrated to be competitive when compared to the most current classifier PhyloPythia and has the advantage that it can be locally installed and the reference set can be kept up-to-date.
Naryttza N. Diaz, Lutz Krause, Alexander Goesmann, Karsten Niehaus, Tim W. Nattkemper
BMC Bioinform.3
2009 EMMA 2 - A MAGE-compliant system for the collaborative analysis and integration of microarray data
abstract
BACKGROUND: Understanding transcriptional regulation by genome-wide microarray studies can contribute to unravel complex relationships between genes. Attempts to standardize the annotation of microarray data include the Minimum Information About a Microarray Experiment (MIAME) recommendations, the MAGE-ML format for data interchange, and the use of controlled vocabularies or ontologies. The existing software systems for microarray data analysis implement the mentioned standards only partially and are often hard to use and extend. Integration of genomic annotation data and other sources of external knowledge using open standards is therefore a key requirement for future integrated analysis systems. RESULTS: The EMMA 2 software has been designed to resolve shortcomings with respect to full MAGE-ML and ontology support and makes use of modern data integration techniques. We present a software system that features comprehensive data analysis functions for spotted arrays, and for the most common synthesized oligo arrays such as Agilent, Affymetrix and NimbleGen. The system is based on the full MAGE object model. Analysis functionality is based on R and Bioconductor packages and can make use of a compute cluster for distributed services. CONCLUSION: Our model-driven approach for automatically implementing a full MAGE object model provides high flexibility and compatibility. Data integration via SOAP-based web-services is advantageous in a distributed client-server environment as the collaborative analysis of microarray data is gaining more and more relevance in international research consortia. The adequacy of the EMMA 2 software design and implementation has been proven by its application in many distributed functional genomics projects. Its scalability makes the current architecture suited for extensions towards future transcriptomics methods based on high-throughput sequencing approaches which have much higher computational requirements than microarrays.
Michael Dondrup, Stefan P. Albaum, Thasso Griebel, Kolja Henckel, Sebastian Jünemann, Tim Kahlke, Christiane K. Kleindt, Helge Küster, Burkhard Linke, Dominik Mertens, Virginie Mittard-Runte, Heiko Neuweger, Kai J. Runte, Andreas Tauch, Felix Tille, Alfred Pühler, Alexander Goesmann
BMC Bioinform.17
2009 WebCARMA: a web application for the functional and taxonomic classification of unassembled metagenomic reads
abstract
BACKGROUND: Metagenomics is a new field of research on natural microbial communities. High-throughput sequencing techniques like 454 or Solexa-Illumina promise new possibilities as they are able to produce huge amounts of data in much shorter time and with less efforts and costs than the traditional Sanger technique. But the data produced comes in even shorter reads (35-100 basepairs with Illumina, 100-500 basepairs with 454-sequencing). CARMA is a new software pipeline for the characterisation of species composition and the genetic potential of microbial samples using short, unassembled reads. RESULTS: In this paper, we introduce WebCARMA, a refined version of CARMA available as a web application for the taxonomic and functional classification of unassembled (ultra-)short reads from metagenomic communities. In addition, we have analysed the applicability of ultra-short reads in metagenomics. CONCLUSIONS: We show that unassembled reads as short as 35 bp can be used for the taxonomic classification of a metagenome. The web application is freely available at http://webcarma.cebitec.uni-bielefeld.de.
Wolfgang Gerlach, Sebastian Jünemann, Felix Tille, Alexander Goesmann, Jens Stoye
BMC Bioinform.4
2008 MeltDB: a software platform for the analysis and integration of metabolomics experiment data
abstract
MOTIVATION: The recent advances in metabolomics have created the potential to measure the levels of hundreds of metabolites which are the end products of cellular regulatory processes. The automation of the sample acquisition and subsequent analysis in high-throughput instruments that are capable of measuring metabolites is posing a challenge on the necessary systematic storage and computational processing of the experimental datasets. Whereas a multitude of specialized software systems for individual instruments and preprocessing methods exists, there is clearly a need for a free and platform-independent system that allows the standardized and integrated storage and analysis of data obtained from metabolomics experiments. Currently there exists no such system that on the one hand supports preprocessing of raw datasets but also allows to visualize and integrate the results of higher level statistical analyses within a functional genomics context. RESULTS: To facilitate the systematic storage, analysis and integration of metabolomics experiments, we have implemented MeltDB, a web-based software platform for the analysis and annotation of datasets from metabolomics experiments. MeltDB supports open file formats (netCDF, mzXML, mzDATA) and facilitates the integration and evaluation of existing preprocessing methods. The system provides researchers with means to consistently describe and store their experimental datasets. Comprehensive analysis and visualization features of metabolomics datasets are offered to the community through a web-based user interface. The system covers the process from raw data to the visualization of results in a knowledge-based background and is integrated into the context of existing software platforms of genomics and transcriptomics at Bielefeld University. We demonstrate the potential of MeltDB by means of a sample experiment where we dissect the influence of three different carbon sources on the gram-negative bacterium Xanthomonas campestris pv. campestris on the level of measured metabolites. Experimental data are stored, analyzed and annotated within MeltDB and accessible via the public MeltDB web server. AVAILABILITY: The system is publicly available at http://meltdb.cebitec.uni-bielefeld.de.
Heiko Neuweger, Stefan P. Albaum, Michael Dondrup, Marcus Persicke, Tony Watt, Karsten Niehaus, Jens Stoye, Alexander Goesmann
Bioinform.8
2005 BACCardI-a tool for the validation of genomic assemblies, assisting genome finishing and intergenome comparison
abstract
SUMMARY: We provide the graphical tool BACCardI for the construction of virtual clone maps from standard assembler output files or BLAST based sequence comparisons. This new tool has been applied to numerous genome projects to solve various problems including (a) validation of whole genome shotgun assemblies, (b) support for contig ordering in the finishing phase of a genome project, and (c) intergenome comparison between related strains when only one of the strains has been sequenced and a large insert library is available for the other. The BACCardI software can seamlessly interact with various sequence assembly packages. MOTIVATION: Genomic assemblies generated from sequence information need to be validated by independent methods such as physical maps. The time-consuming task of building physical maps can be circumvented by virtual clone maps derived from read pair information of large insert libraries.
Daniela Bartels, Sebastian Kespohl, Stefan P. Albaum, Tanja Drüke, Alexander Goesmann, Julia Herold, Olaf Kaiser, Alfred Pühler, Friedhelm Pfeiffer, Günter Raddatz, Jens Stoye, Folker Meyer, Stephan C. Schuster
Bioinform.5
2004 Development of joint application strategies for two microbial gene finders
abstract
MOTIVATION: As a starting point in annotation of bacterial genomes, gene finding programs are used for the prediction of functional elements in the DNA sequence. Due to the faster pace and increasing number of genome projects currently underway, it is becoming especially important to have performant methods for this task. RESULTS: This study describes the development of joint application strategies that combine the strengths of two microbial gene finders to improve the overall gene finding performance. Critica is very specific in the detection of similarity-supported genes as it uses a comparative sequence analysis-based approach. Glimmer employs a very sophisticated model of genomic sequence properties and is sensitive also in the detection of organism-specific genes. Based on a data set of 113 microbial genome sequences, we optimized a combined application approach using different parameters with relevance to the gene finding problem. This results in a significant improvement in specificity while there is similarity in sensitivity to Glimmer. The improvement is especially pronounced for GC rich genomes. The method is currently being applied for the annotation of several microbial genomes. AVAILABILITY: The methods described have been implemented within the gene prediction component of the GenDB genome annotation system.
Alice C. McHardy, Alexander Goesmann, Alfred Pühler, Folker Meyer
Bioinform.2
2002 PathFinder: reconstruction and dynamic visualization of metabolic pathways
abstract
Abstract Motivation: Beyond methods for a gene-wise annotation and analysis of sequenced genomes new automated methods for functional analysis on a higher level are needed. The identification of realized metabolic pathways provides valuable information on gene expression and regulation. Detection of incomplete pathways helps to improve a constantly evolving genome annotation or discover alternative biochemical pathways. To utilize automated genome analysis on the level of metabolic pathways new methods for the dynamic representation and visualization of pathways are needed. Results: PathFinder is a tool for the dynamic visualization of metabolic pathways based on annotation data. Pathways are represented as directed acyclic graphs, graph layout algorithms accomplish the dynamic drawing and visualization of the metabolic maps. A more detailed analysis of the input data on the level of biochemical pathways helps to identify genes and detect improper parts of annotations. As an Relational Database Management System (RDBMS) based internet application PathFinder reads a list of EC-numbers or a given annotation in EMBL- or Genbank-format and dynamically generates pathway graphs. Availability: The software PathFinder will be made available on the Bielefeld Bioinformatics WebServer under the following URL: http://bibiserv.TechFak.Uni-Bielefeld.DE/pathfinder/. The source code is available upon request and will eventually be released under GPL. Contact: [email protected]
Alexander Goesmann, Martin Haubrock, Folker Meyer, Jörn Kalinowski, Robert Giegerich
Bioinform.1