Andreas Hildebrandt 0001

dblp:39/2195 · DBLP profile ↗
← Back
27ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0003-2180-6516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Cross-domain transfer learning from peptides to metabolites using a multi-property fine-tuned LLM
abstract
MOTIVATION: Accurate liquid chromatography retention time (RT) prediction is a critical component of compound identification in metabolomics and lipidomics. However, existing RT prediction approaches are often limited by the scarcity of experimental RT measurements for many molecular classes, restricting model generalization and the construction of comprehensive RT libraries. Transfer learning from data-rich chemical domains offers a potential strategy to overcome these limitations, but its effectiveness for metabolite RT prediction remains insufficiently explored. RESULTS: We developed a transfer learning framework based on ChemBERTa that leverages large peptide datasets to improve metabolite RT prediction under data-sparse conditions. A peptide-pretrained model was trained using a multi-task objective that jointly predicted RT and seven RDKit-derived molecular descriptors. Compared with an RT-only model, the multi-task approach learned more robust chemical representations and demonstrated superior generalization to metabolites, achieving a median test R² of 0.842 versus 0.820. When transferred to metabolite RT prediction, the multi-task pretrained model substantially outperformed models trained from scratch at low-data regimes. Using only 3% of metabolite training data (2129 compounds), transfer learning achieved a median test R² of 0.322 compared with 0.216 for the baseline model, while reducing MAE from 131.7 to 114.9. Significant improvements were also observed at 5% and 10% training fractions, with benefits gradually diminishing as larger metabolite datasets became available. In contrast, a peptide-pretrained single-task RT model showed performance comparable to the baseline, indicating that the observed gains arise primarily from multi-task molecular property learning rather than peptide pretraining alone. These findings demonstrate that multi-task transfer learning provides an effective and scalable strategy for improving RT prediction in metabolomics, particularly when experimental training data are limited. AVAILABILITY: Freely available on https://github.com/uchealex/CHEMBEDDING.
Uchenna Alex Anyaegbunam, David Teschner, Thierry Schmidlin, Andreas Hildebrandt 0001, Johannes U. Mayer, Maximilian Sprang, Miguel A. Andrade-Navarro
Bioinform.4
2023 Ionmob: a Python package for prediction of peptide collisional cross-section values
abstract
MOTIVATION: Including ion mobility separation (IMS) into mass spectrometry proteomics experiments is useful to improve coverage and throughput. Many IMS devices enable linking experimentally derived mobility of an ion to its collisional cross-section (CCS), a highly reproducible physicochemical property dependent on the ion's mass, charge and conformation in the gas phase. Thus, known peptide ion mobilities can be used to tailor acquisition methods or to refine database search results. The large space of potential peptide sequences, driven also by posttranslational modifications of amino acids, motivates an in silico predictor for peptide CCS. Recent studies explored the general performance of varying machine-learning techniques, however, the workflow engineering part was of secondary importance. For the sake of applicability, such a tool should be generic, data driven, and offer the possibility to be easily adapted to individual workflows for experimental design and data processing. RESULTS: We created ionmob, a Python-based framework for data preparation, training, and prediction of collisional cross-section values of peptides. It is easily customizable and includes a set of pretrained, ready-to-use models and preprocessing routines for training and inference. Using a set of ≈21 000 unique phosphorylated peptides and ≈17 000 MHC ligand sequences and charge state pairs, we expand upon the space of peptides that can be integrated into CCS prediction. Lastly, we investigate the applicability of in silico predicted CCS to increase confidence in identified peptides by applying methods of re-scoring and demonstrate that predicted CCS values complement existing predictors for that task. AVAILABILITY AND IMPLEMENTATION: The Python package is available at github: https://github.com/theGreatHerrLebert/ionmob.
David Teschner, David Gomez-Zepeda, Arthur Declercq, Mateusz K. Lacki, Seymen Avci, Konstantin Bob, Ute Distler, Thomas Michna, Lennart Martens, Stefan Tenzer, Andreas Hildebrandt 0001
Bioinform.11
2022 Locality-sensitive hashing enables efficient and scalable signal classification in high-throughput mass spectrometry raw data
abstract
BACKGROUND: Mass spectrometry is an important experimental technique in the field of proteomics. However, analysis of certain mass spectrometry data faces a combination of two challenges: first, even a single experiment produces a large amount of multi-dimensional raw data and, second, signals of interest are not single peaks but patterns of peaks that span along the different dimensions. The rapidly growing amount of mass spectrometry data increases the demand for scalable solutions. Furthermore, existing approaches for signal detection usually rely on strong assumptions concerning the signals properties. RESULTS: In this study, it is shown that locality-sensitive hashing enables signal classification in mass spectrometry raw data at scale. Through appropriate choice of algorithm parameters it is possible to balance false-positive and false-negative rates. On synthetic data, a superior performance compared to an intensity thresholding approach was achieved. Real data could be strongly reduced without losing relevant information. Our implementation scaled out up to 32 threads and supports acceleration by GPUs. CONCLUSIONS: Locality-sensitive hashing is a desirable approach for signal classification in mass spectrometry raw data. AVAILABILITY: Generated data and code are available at https://github.com/hildebrandtlab/mzBucket . Raw data is available at https://zenodo.org/record/5036526 .
Konstantin Bob, David Teschner, Thomas Kemmer, David Gomez-Zepeda, Stefan Tenzer, Bertil Schmidt, Andreas Hildebrandt 0001
BMC Bioinform.7
2021 CARE: context-aware sequencing read error correction
abstract
MOTIVATION: Error correction is a fundamental pre-processing step in many Next-Generation Sequencing (NGS) pipelines, in particular for de novo genome assembly. However, existing error correction methods either suffer from high false-positive rates since they break reads into independent k-mers or do not scale efficiently to large amounts of sequencing reads and complex genomes. RESULTS: We present CARE-an alignment-based scalable error correction algorithm for Illumina data using the concept of minhashing. Minhashing allows for efficient similarity search within large sequencing read collections which enables fast computation of high-quality multiple alignments. Sequencing errors are corrected by detailed inspection of the corresponding alignments. Our performance evaluation shows that CARE generates significantly fewer false-positive corrections than state-of-the-art tools (Musket, SGA, BFC, Lighter, Bcool, Karect) while maintaining a competitive number of true positives. When used prior to assembly it can achieve superior de novo assembly results for a number of real datasets. CARE is also the first multiple sequence alignment-based error corrector that is able to process a human genome Illumina NGS dataset in only 4 h on a single workstation using GPU acceleration. AVAILABILITYAND IMPLEMENTATION: CARE is open-source software written in C++ (CPU version) and in CUDA/C++ (GPU version). It is licensed under GPLv3 and can be downloaded at https://github.com/fkallen/CARE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Felix Kallenborn, Andreas Hildebrandt 0001, Bertil Schmidt
Bioinform.2
2020 AnySeq: A High Performance Sequence Alignment Library based on Partial Evaluation
abstract
Sequence alignments are fundamental to bioinformatics which has resulted in a variety of optimized implementations. Unfortunately, the vast majority of them are hand-tuned and specific to certain architectures and execution models. This not only makes them challenging to understand and extend, but also difficult to port to other platforms. We present AnySeq - a novel library for computing different types of pairwise alignments of DNA sequences. Our approach combines high performance with an intuitively understandable implementation, which is achieved through the concept of partial evaluation. Using the AnyDSL compiler framework, AnySeq enables the compilation of algorithmic variants that are highly optimized for specific usage scenarios and hardware targets with a single, uniform codebase. The resulting domain-specific library thus allows the variation of alignment parameters (such as alignment type, scoring scheme, and traceback vs.~plain score) by simple function composition rather than metaprogramming techniques which are often hard to understand. Our implementation supports multithreading and SIMD vectorization on CPUs, CUDA-enabled GPUs, and FPGAs. AnySeq is at most 7% slower and in many cases faster (up to 12%) than state-of-the art manually optimized alignment libraries on CPUs (SeqAn) and on GPUs (NVBio).
André Müller, Bertil Schmidt, Andreas Hildebrandt 0001, Richard Membarth, Roland Leißa, Matthis Kruse, Sebastian Hack
IPDPS3
2020 A big data approach to metagenomics for all-food-sequencing
abstract
BACKGROUND: All-Food-Sequencing (AFS) is an untargeted metagenomic sequencing method that allows for the detection and quantification of food ingredients including animals, plants, and microbiota. While this approach avoids some of the shortcomings of targeted PCR-based methods, it requires the comparison of sequence reads to large collections of reference genomes. The steadily increasing amount of available reference genomes establishes the need for efficient big data approaches. RESULTS: We introduce an alignment-free k-mer based method for detection and quantification of species composition in food and other complex biological matters. It is orders-of-magnitude faster than our previous alignment-based AFS pipeline. In comparison to the established tools CLARK, Kraken2, and Kraken2+Bracken it is superior in terms of false-positive rate and quantification accuracy. Furthermore, the usage of an efficient database partitioning scheme allows for the processing of massive collections of reference genomes with reduced memory requirements on a workstation (AFS-MetaCache) or on a Spark-based compute cluster (MetaCacheSpark). CONCLUSIONS: We present a fast yet accurate screening method for whole genome shotgun sequencing-based biosurveillance applications such as food testing. By relying on a big data approach it can scale efficiently towards large-scale collections of complex eukaryotic and bacterial reference genomes. AFS-MetaCache and MetaCacheSpark are suitable tools for broad-scale metagenomic screening applications. They are available at https://muellan.github.io/metacache/afs.html (C++ version for a workstation) and https://github.com/jmabuin/MetaCacheSpark (Spark version for big data clusters).
Robin Kobus, José Manuel Abuín, André Müller, Sören Lukas Hellmann, Juan Carlos Pichel, Tomás F. Pena, Andreas Hildebrandt 0001, Thomas Hankeln, Bertil Schmidt
BMC Bioinform.7
2017 MetaCache: context-aware classification of metagenomic reads using minhashing
abstract
MOTIVATION: Metagenomic shotgun sequencing studies are becoming increasingly popular with prominent examples including the sequencing of human microbiomes and diverse environments. A fundamental computational problem in this context is read classification, i.e. the assignment of each read to a taxonomic label. Due to the large number of reads produced by modern high-throughput sequencing technologies and the rapidly increasing number of available reference genomes corresponding software tools suffer from either long runtimes, large memory requirements or low accuracy. RESULTS: We introduce MetaCache-a novel software for read classification using the big data technique minhashing. Our approach performs context-aware classification of reads by computing representative subsamples of k-mers within both, probed reads and locally constrained regions of the reference genomes. As a result, MetaCache consumes significantly less memory compared to the state-of-the-art read classifiers Kraken and CLARK while achieving highly competitive sensitivity and precision at comparable speed. For example, using NCBI RefSeq draft and completed genomes with a total length of around 140 billion bases as reference, MetaCache's database consumes only 62 GB of memory while both Kraken and CLARK fail to construct their respective databases on a workstation with 512 GB RAM. Our experimental results further show that classification accuracy continuously improves when increasing the amount of utilized reference genome data. AVAILABILITY AND IMPLEMENTATION: MetaCache is open source software written in C ++ and can be downloaded at http://github.com/muellan/metacache. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
André Müller, Christian Hundt 0002, Andreas Hildebrandt 0001, Thomas Hankeln, Bertil Schmidt
Bioinform.3
2016 rapidGSEA: Speeding up gene set enrichment analysis on multi-core CPUs and CUDA-enabled GPUs
abstract
BACKGROUND: Gene Set Enrichment Analysis (GSEA) is a popular method to reveal significant dependencies between predefined sets of gene symbols and observed phenotypes by evaluating the deviation of gene expression values between cases and controls. An established measure of inter-class deviation, the enrichment score, is usually computed using a weighted running sum statistic over the whole set of gene symbols. Due to the lack of analytic expressions the significance of enrichment scores is determined using a non-parametric estimation of their null distribution by permuting the phenotype labels of the probed patients. Accordingly, GSEA is a time-consuming task due to the large number of required permutations to accurately estimate the nominal p-value - a circumstance that is even more pronounced during multiple hypothesis testing since its estimate is lower-bounded by the inverse number of samples in permutation space. RESULTS: We present rapidGSEA - a software suite consisting of two tools for facilitating permutation-based GSEA: cudaGSEA and ompGSEA. cudaGSEA is a CUDA-accelerated tool using fine-grained parallelization schemes on massively parallel architectures while ompGSEA is a coarse-grained multi-threaded tool for multi-core CPUs. Nominal p-value estimation of 4,725 gene sets on a data set consisting of 20,639 unique gene symbols and 200 patients (183 cases + 17 controls) each probing one million permutations takes 19 hours on a Xeon CPU and less than one hour on a GeForce Titan X GPU while the established GSEA tool from the Broad Institute (broadGSEA) takes roughly 13 days. CONCLUSION: cudaGSEA outperforms broadGSEA by around two orders-of-magnitude on a single Tesla K40c or GeForce Titan X GPU. ompGSEA provides around one order-of-magnitude speedup to broadGSEA on a standard Xeon CPU. The rapidGSEA suite is open-source software and can be downloaded at https://github.com/gravitino/cudaGSEA as standalone application or package for the R framework.
Christian Hundt 0002, Andreas Hildebrandt 0001, Bertil Schmidt
BMC Bioinform.2
2015 ballaxy: web services for structural bioinformatics
abstract
MOTIVATION: Web-based workflow systems have gained considerable momentum in sequence-oriented bioinformatics. In structural bioinformatics, however, such systems are still relatively rare; while commercial stand-alone workflow applications are common in the pharmaceutical industry, academic researchers often still rely on command-line scripting to glue individual tools together. RESULTS: In this work, we address the problem of building a web-based system for workflows in structural bioinformatics. For the underlying molecular modelling engine, we opted for the BALL framework because of its extensive and well-tested functionality in the field of structural bioinformatics. The large number of molecular data structures and algorithms implemented in BALL allows for elegant and sophisticated development of new approaches in the field. We hence connected the versatile BALL library and its visualization and editing front end BALLView with the Galaxy workflow framework. The result, which we call ballaxy, enables the user to simply and intuitively create sophisticated pipelines for applications in structure-based computational biology, integrated into a standard tool for molecular modelling. AVAILABILITY AND IMPLEMENTATION: ballaxy consists of three parts: some minor modifications to the Galaxy system, a collection of tools and an integration into the BALL framework and the BALLView application for molecular modelling. Modifications to Galaxy will be submitted to the Galaxy project, and the BALL and BALLView integrations will be integrated in the next major BALL release. After acceptance of the modifications into the Galaxy project, we will publish all ballaxy tools via the Galaxy toolshed. In the meantime, all three components are available from http://www.ball-project.org/ballaxy. Also, docker images for ballaxy are available at https://registry.hub.docker.com/u/anhi/ballaxy/dockerfile/. ballaxy is licensed under the terms of the GPL.
Anna Katharina Hildebrandt, Daniel Stöckel, Nina M. Fischer, Luis de la Garza, Jens Krüger 0002, Stefan Nickels, Marc Röttig, Charlotta Schärfe, Marcel Schumann, Philipp Thiel, Hans-Peter Lenhof, Oliver Kohlbacher, Andreas Hildebrandt 0001
Bioinform.13
2014 Algorithms for the Maximum Weight Connected k -Induced Subgraph Problem
Ernst Althaus, Markus Blumenstock, Alexej Disterhoft, Andreas Hildebrandt 0001, Markus Krupp
COCOA4
2014 Parallelized Clustering of Protein Structures on CUDA-Enabled GPUs
abstract
Estimation of the pose in which two given molecules might bind together to form a potential complex is a crucial task in structural biology. To solve this so-called "docking problem", most algorithms initially generate large numbers of candidate poses (or decoys) which are then clustered to allow for subsequent computationally expensive evaluations of reasonable representatives. Since the number of such candidates ranges from thousands to millions, performing the clustering on standard CPUs is highly time consuming. In this paper we analyze and evaluate different approaches to parallelize the nearest neighbor chain algorithm to perform hierarchical Ward clustering of protein structures using both atom-based root mean square deviation (RMSD) and rigid-based RMSD molecular distances on a GPU. This leads to a speedup of around one order-of-magnitude of our CUDA implementation on a GeForce Titan GPU compared to a multi-threaded CPU implementation on a Core-i7 2700.
Hoang-Vu Dang, Bertil Schmidt, Andreas Hildebrandt 0001, Anna Katharina Hildebrandt
PDP3
2014 SKINK: a web server for string kernel based kink prediction in α-helices
abstract
MOTIVATION: The reasons for distortions from optimal α-helical geometry are widely unknown, but their influences on structural changes of proteins are significant. Hence, their prediction is a crucial problem in structural bioinformatics. Here, we present a new web server, called SKINK, for string kernel based kink prediction. Extending our previous study, we also annotate the most probable kink position in a given α-helix sequence. AVAILABILITY AND IMPLEMENTATION: The SKINK web server is freely accessible at http://biows-inf.zdv.uni-mainz.de/skink. Moreover, SKINK is a module of the BALL software, also freely available at www.ballview.org.
Tim Seifert, Andreas Lund 0001, Benny Kneissl, Sabine C. Mueller, Christofer S. Tautermann, Andreas Hildebrandt 0001
Bioinform.6
2013 NightShift: NMR Shift Inference by General Hybrid Model Training - a Framework for NMR Chemical Shift Prediction
abstract
BACKGROUND: NMR chemical shift prediction plays an important role in various applications in computational biology. Among others, structure determination, structure optimization, and the scoring of docking results can profit from efficient and accurate chemical shift estimation from a three-dimensional model.A variety of NMR chemical shift prediction approaches have been presented in the past, but nearly all of these rely on laborious manual data set preparation and the training itself is not automatized, making retraining the model, e.g., if new data is made available, or testing new models a time-consuming manual chore. RESULTS: In this work, we present the framework NightShift (NMR Shift Inference by General Hybrid Model Training), which enables automated data set generation as well as model training and evaluation of protein NMR chemical shift prediction.In addition to this main result - the NightShift framework itself - we describe the resulting, automatically generated, data set and, as a proof-of-concept, a random forest model called Spinster that was built using the pipeline. CONCLUSION: By demonstrating that the performance of the automatically generated predictors is at least en par with the state of the art, we conclude that automated data set and predictor generation is well-suited for the design of NMR chemical shift estimators.The framework can be downloaded from https://bitbucket.org/akdehof/nightshift. It requires the open source Biochemical Algorithms Library (BALL), and is available under the conditions of the GNU Lesser General Public License (LGPL). We additionally offer a browser-based user interface to our NightShift instance employing the Galaxy framework via https://ballaxy.bioinf.uni-sb.de/.
Anna Katharina Hildebrandt, Simon Loew, Hans-Peter Lenhof, Andreas Hildebrandt 0001
BMC Bioinform.4
2012 ProteinScanAR - An Augmented Reality Web Application for High School Education in Biomolecular Life Sciences
abstract
Understanding protein structures is a crucial step in creating molecular insight for researchers as well as students and pupils. The enormous scaling gap between an atomic point of view and objects in daily life hampers developing an intuitive relation between them. Especially for high school students, it can be difficult to understand the spatial relations of a protein structure. Due to lack of direct imaging techniques, molecules can only be explored by studying abstract molecular models. Here, the use of Augmented reality (AR) techniques has proven to strongly improve structural perception. In this work we present ProteinScanAR, an augmented reality framework for biomolecular education that allows connecting virtual and real worlds intuitively, and thus enables focusing on the scientific or educational content. Special attention was taken to guarantee implementational and technical requirements as general and simple as possible to alleviate application in nonexpert computer settings. The ProteinScanAR framework is freely available under the GNU Public License (GPL).
Stefan Nickels, Hienke Sminia, Sabine C. Mueller, Bas Kools, Anna Katharina Hildebrandt, Hans-Peter Lenhof, Andreas Hildebrandt 0001
IV7
2012 A dynamic program analysis to find floating-point accuracy problems
abstract
Programs using floating-point arithmetic are prone to accuracy problems caused by rounding and catastrophic cancellation. These phenomena provoke bugs that are notoriously hard to track down: the program does not necessarily crash and the results are not necessarily obviously wrong, but often subtly inaccurate. Further use of these values can lead to catastrophic errors.
Florian Benz, Andreas Hildebrandt 0001, Sebastian Hack
PLDI2
2012 Isotope pattern deconvolution for peptide mass spectrometry by non-negative least squares/least absolute deviation template matching
abstract
BACKGROUND: The robust identification of isotope patterns originating from peptides being analyzed through mass spectrometry (MS) is often significantly hampered by noise artifacts and the interference of overlapping patterns arising e.g. from post-translational modifications. As the classification of the recorded data points into either 'noise' or 'signal' lies at the very root of essentially every proteomic application, the quality of the automated processing of mass spectra can significantly influence the way the data might be interpreted within a given biological context. RESULTS: We propose non-negative least squares/non-negative least absolute deviation regression to fit a raw spectrum by templates imitating isotope patterns. In a carefully designed validation scheme, we show that the method exhibits excellent performance in pattern picking. It is demonstrated that the method is able to disentangle complicated overlaps of patterns. CONCLUSIONS: We find that regularization is not necessary to prevent overfitting and that thresholding is an effective and user-friendly way to perform feature selection. The proposed method avoids problems inherent in regularization-based approaches, comes with a set of well-interpretable parameters whose default configuration is shown to generalize well without the need for fine-tuning, and is applicable to spectra of different platforms. The R package IPPD implements the method and is available from the Bioconductor platform (http://bioconductor.fhcrc.org/help/bioc-views/devel/bioc/html/IPPD.html).
Martin Slawski, Rene Hussong, Andreas Tholey, Thomas Jakoby, Barbara Gregorius, Andreas Hildebrandt 0001, Matthias Hein 0001
BMC Bioinform.6
2011 A fast solver for nonlocal electrostatic theory in biomolecular science and engineering
abstract
Biological molecules perform their functions surrounded by water and mobile ions, which strongly influence molecular structure and behavior. The electrostatic interactions between a molecule and solvent are particularly difficult to model theoretically, due to the forces' long range and the collective response of many thousands of solvent molecules. The dominant modeling approaches represent the two extremes of the trade-off between molecular realism and computational efficiency: all-atom molecular dynamics in explicit solvent, and macroscopic continuum theory (the Poisson or Poisson--Boltzmann equation). We present the first fast-solver implementation of an advanced nonlocal continuum theory that combines key advantages of both approaches. In particular, molecular realism is included by limiting solvent dielectric response on short length scales, using a model for nonlocal dielectric response allows the resulting problem (a linear integro-differential Poisson equation) to be reformulated as a system of coupled boundary-integral equations using double reciprocity. Whereas previous studies using the nonlocal theory had been limited to small model problems, owing to computational cost, our work opens the door to studying much larger problems including rational drug design, protein engineering, and nanofluidics.
Jaydeep P. Bardhan, Andreas Hildebrandt 0001
DAC2
2011 Automated bond order assignment as an optimization problem
abstract
MOTIVATION: Numerous applications in Computational Biology process molecular structures and hence strongly rely not only on correct atomic coordinates but also on correct bond order information. For proteins and nucleic acids, bond orders can be easily deduced but this does not hold for other types of molecules like ligands. For ligands, bond order information is not always provided in molecular databases and thus a variety of approaches tackling this problem have been developed. In this work, we extend an ansatz proposed by Wang et al. that assigns connectivity-based penalty scores and tries to heuristically approximate its optimum. In this work, we present three efficient and exact solvers for the problem replacing the heuristic approximation scheme of the original approach: an A*, an ILP and an fixed-parameter approach (FPT) approach. RESULTS: We implemented and evaluated the original implementation, our A*, ILP and FPT formulation on the MMFF94 validation suite and the KEGG Drug database. We show the benefit of computing exact solutions of the penalty minimization problem and the additional gain when computing all optimal (or even suboptimal) solutions. We close with a detailed comparison of our methods. AVAILABILITY: The A* and ILP solution are integrated into the open-source C++ LGPL library BALL and the molecular visualization and modelling tool BALLView and can be downloaded from our homepage www.ball-project.org. The FPT implementation can be downloaded from http://bio.informatik.uni-jena.de/software/.
Anna Katharina Hildebrandt, Alexander Rurainski, Quang Bao Anh Bui, Sebastian Böcker, Hans-Peter Lenhof, Andreas Hildebrandt 0001
Bioinform.6
2010 Real-Time Ray Tracing of Complex Molecular Scenes
abstract
Molecular visualization is one of the cornerstones in structural bioinformatics and related fields. Today, rasterization is typically used for the interactive display of molecular scenes, while ray tracing aims at generating high-quality images, taking typically minutes to hours to generate and requiring the usage of an external off-line program. Recently, real-time ray tracing evolved to combine the interactivity of rasterization-based approaches with the superb image quality of ray tracing techniques. We demonstrate how real-time ray tracing integrated into a molecular modelling and visualization tool allows for better understanding of the structural arrangement of biomolecules and natural creation of publication-quality images in real-time. However, unlike most approaches, our technique naturally integrates into the full-featured molecular modelling and visualization tool BALL View, seamlessly extending a standard workflow with interactive high-quality rendering.
Lukas Marsalek, Anna Katharina Hildebrandt, Iliyan Georgiev, Hans-Peter Lenhof, Philipp Slusallek, Andreas Hildebrandt 0001
IV6
2010 BALL - biochemical algorithms library 1.3
abstract
BACKGROUND: The Biochemical Algorithms Library (BALL) is a comprehensive rapid application development framework for structural bioinformatics. It provides an extensive C++ class library of data structures and algorithms for molecular modeling and structural bioinformatics. Using BALL as a programming toolbox does not only allow to greatly reduce application development times but also helps in ensuring stability and correctness by avoiding the error-prone reimplementation of complex algorithms and replacing them with calls into the library that has been well-tested by a large number of developers. In the ten years since its original publication, BALL has seen a substantial increase in functionality and numerous other improvements. RESULTS: Here, we discuss BALL's current functionality and highlight the key additions and improvements: support for additional file formats, molecular edit-functionality, new molecular mechanics force fields, novel energy minimization techniques, docking algorithms, and support for cheminformatics. CONCLUSIONS: BALL is available for all major operating systems, including Linux, Windows, and MacOS X. It is available free of charge under the Lesser GNU Public License (LPGL). Parts of the code are distributed under the GNU Public License (GPL). BALL is available as source code and binary packages from the project web site at http://www.ball-project.org. Recently, it has been accepted into the debian project; integration into further distributions is currently pursued.
Andreas Hildebrandt 0001, Anna Katharina Hildebrandt, Alexander Rurainski, Andreas Bertsch, Marcel Schumann, Nora C. Toussaint, Andreas Moll, Daniel Stöckel, Stefan Nickels, Sabine C. Mueller, Hans-Peter Lenhof, Oliver Kohlbacher
BMC Bioinform.1
2009 Highly accelerated feature detection in proteomics data sets using modern graphics processing units
abstract
MOTIVATION: Mass spectrometry (MS) is one of the most important techniques for high-throughput analysis in proteomics research. Due to the large number of different proteins and their post-translationally modified variants, the amount of data generated by a single wet-lab MS experiment can easily exceed several gigabytes. Hence, the time necessary to analyze and interpret the measured data is often significantly larger than the time spent on sample preparation and the wet-lab experiment itself. Since the automated analysis of this data is hampered by noise and baseline artifacts, more sophisticated computational techniques are required to handle the recorded mass spectra. Obviously, there is a clear tradeoff between performance and quality of the analysis, which is currently one of the most challenging problems in computational proteomics. RESULTS: Using modern graphics processing units (GPUs), we implemented a feature finding algorithm based on a hand-tailored adaptive wavelet transform that drastically reduces the computation time. A further speedup can be achieved exploiting the multi-core architecture of current computing devices, which leads to up to an approximately 200-fold speed-up in our computational experiments. In addition, we will demonstrate that several approximations necessary on the CPU to keep run times bearable, become obsolete on the GPU, yielding not only faster, but also improved results. AVAILABILITY: An open source implementation of the CUDA-based algorithm is available via the software framework OpenMS (http://www.openms.de). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rene Hussong, Barbara Gregorius, Andreas Tholey, Andreas Hildebrandt 0001
Bioinform.4
2008 OpenMS - An open-source software framework for mass spectrometry
abstract
BACKGROUND: Mass spectrometry is an essential analytical technique for high-throughput analysis in proteomics and metabolomics. The development of new separation techniques, precise mass analyzers and experimental protocols is a very active field of research. This leads to more complex experimental setups yielding ever increasing amounts of data. Consequently, analysis of the data is currently often the bottleneck for experimental studies. Although software tools for many data analysis tasks are available today, they are often hard to combine with each other or not flexible enough to allow for rapid prototyping of a new analysis workflow. RESULTS: We present OpenMS, a software framework for rapid application development in mass spectrometry. OpenMS has been designed to be portable, easy-to-use and robust while offering a rich functionality ranging from basic data structures to sophisticated algorithms for data analysis. This has already been demonstrated in several studies. CONCLUSION: OpenMS is available under the Lesser GNU Public License (LGPL) from the project website at http://www.openms.de.
Marc Sturm, Andreas Bertsch, Clemens Gröpl, Andreas Hildebrandt 0001, Rene Hussong, Eva Lange, Nico Pfeifer, Ole Schulz-Trieglaff, Alexandra Zerck, Knut Reinert, Oliver Kohlbacher
BMC Bioinform.4
2007 A Fast and Accurate Algorithm for the Quantification of Peptides from Mass Spectrometry Data
Ole Schulz-Trieglaff, Rene Hussong, Clemens Gröpl, Andreas Hildebrandt 0001, Knut Reinert
RECOMB4
2007 Electrostatic potentials of proteins in water: a structured continuum approach
abstract
Electrostatic interactions play a crucial role in many biomolecular processes, including molecular recognition and binding. Biomolecular electrostatics is modulated to a large extent by the water surrounding the molecules. Here, we present a novel approach to the computation of electrostatic potentials which allows the inclusion of water structure into the classical theory of continuum electrostatics. Based on our recent purely differential formulation of nonlocal electrostatics [Hildebrandt, et al. (2004) Phys. Rev. Lett., 93, 108104] we have developed a new algorithm for its efficient numerical solution. The key component of this algorithm is a boundary element solver, having the same computational complexity as established boundary element methods for local continuum electrostatics. This allows, for the first time, the computation of electrostatic potentials and interactions of large biomolecular systems immersed in water including effects of the solvent's structure in a continuum description. We illustrate the applicability of our approach with two examples, the enzymes trypsin and acetylcholinesterase. The approach is applicable to all problems requiring precise prediction of electrostatic interactions in water, such as protein-ligand and protein-protein docking, folding and chromatin regulation. Initial results indicate that this approach may shed new light on biomolecular electrostatics and on aspects of molecular recognition that classical local electrostatics cannot reveal.
Andreas Hildebrandt 0001, Ralf Blossey, Sergej Rjasanow, Oliver Kohlbacher, Hans-Peter Lenhof
Bioinform.1
2006 BALLView: a tool for research and education in molecular modeling
abstract
Abstract Summary: We present BALLView, a molecular viewer and modeling tool. It combines state-of-the-art visualization capabilities with powerful modeling functionality including implementations of force field methods and continuum electrostatics models. BALLView is a versatile and extensible tool for research in structural bioinformatics and molecular modeling. Furthermore, the convenient and intuitive graphical user interface offers novice users direct access to the full functionality, rendering it ideal for teaching. Through an interface to the object-oriented scripting language Python it is easily extensible. Availability: BALLView is an open source software and runs on all major platforms (Windows, MacOS X, Linux and most Unix flavors). It is available free of charge under the GNU Public License at Contact: [email protected]
Andreas Moll, Andreas Hildebrandt 0001, Hans-Peter Lenhof, Oliver Kohlbacher
Bioinform.2
2006 A minimally invasive multiple marker approach allows highly efficient detection of meningioma tumors
abstract
BACKGROUND: The development of effective frameworks that permit an accurate diagnosis of tumors, especially in their early stages, remains a grand challenge in the field of bioinformatics. Our approach uses statistical learning techniques applied to multiple antigen tumor antigen markers utilizing the immune system as a very sensitive marker of molecular pathological processes. For validation purposes we choose the intracranial meningioma tumors as model system since they occur very frequently, are mostly benign, and are genetically stable. RESULTS: A total of 183 blood samples from 93 meningioma patients (WHO stages I-III) and 90 healthy controls were screened for seroreactivity with a set of 57 meningioma-associated antigens. We tested several established statistical learning methods on the resulting reactivity patterns using 10-fold cross validation. The best performance was achieved by Naïve Bayes Classifiers. With this classification method, our framework, called Minimally Invasive Multiple Marker (MIMM) approach, yielded a specificity of 96.2%, a sensitivity of 84.5%, and an accuracy of 90.3%, the respective area under the ROC curve was 0.957. Detailed analysis revealed that prediction performs particularly well on low-grade (WHO I) tumors, consistent with our goal of early stage tumor detection. For these tumors the best classification result with a specificity of 97.5%, a sensitivity of 91.3%, an accuracy of 95.6%, and an area under the ROC curve of 0.971 was achieved using a set of 12 antigen markers only. This antigen set was detected by a subset selection method based on Mutual Information. Remarkably, our study proves that the inclusion of non-specific antigens, detected not only in tumor but also in normal sera, increases the performance significantly, since non-specific antigens contribute additional diagnostic information. CONCLUSION: Our approach offers the possibility to screen members of risk groups as a matter of routine such that tumors hopefully can be diagnosed immediately after their genesis. The early detection will finally result in a higher cure- and lower morbidity-rate.
Andreas Keller, Nicole Ludwig 0001, Nicole Comtesse, Andreas Hildebrandt 0001, Eckart Meese, Hans-Peter Lenhof
BMC Bioinform.4
2001 A NMR-spectra-based scoring function for protein docking
abstract
A well studied problem in the area of Computational Molecular Biology is the so-called Protein-Protein Docking problem (PPD) that can be formulated as follows: Given two proteins A and B that form a protein complex, compute the 3D-structure of the protein complex AB. Protein docking algorithms can be used to study the driving forces and reaction mechanisms of docking processes. They are also able to speed up the lenghty process of experimental structure elucidation of protein complexes by proposing potential structures. In this paper, we are discussing a variant of the PPD-problem where the input consists of the tertiary structures of A and B plus an unassigned 1H-NMR spectrum of the complex AB. We present a new scoring function for evaluating and ranking potential complex structures produced by a docking algorithm. The scoring function computes a “theoretical” 1H-NMR spectrum for each tentative complex structure and subtracts the calculated spectrum from the experimental spectrum. The absolute areas of the difference spectra are then used to rank the potential complex structures. In contrast to formerly published approaches (e.g. Morelli et. al. [38]) we do not use distance constraints (intermolecular NOE constraints). We have tested the approach with the bound conformations of four protein complexes whose three-dimensional structures are stored in the PDB data bank [5] and whose 1H-NMR shift assignments are available from the BMRB database (BioMagResBank [47]).
Oliver Kohlbacher, Andreas Burchardt, Andreas Moll, Andreas Hildebrandt 0001, Peter Bayer, Hans-Peter Lenhof
RECOMB4