VLDB 2026 Research / reviewers in the wild / expert
Rolf Backofen
dblp:b/RolfBackofen
· DBLP profile ↗
95ranked-venue papers
23as first author
16since 2021 · last 2025
0000-0001-8231-3323ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 66 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-authorTheory of computation · 6 · 6 first-authorSoftware engineering, systems software and programming languages · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing dataabstractThere are various dimensionality reduction techniques for visually inspecting dynamical patterns in time-series single-cell RNA-sequencing (scRNA-seq) data. However, the lack of one-to-one correspondence between cells across time points makes it difficult to uniquely uncover temporal structure in a low-dimensional manifold. The use of different techniques may thus lead to discrepancies in the representation of dynamical patterns. However, The extent of these discrepancies remains unclear. To investigate this, we propose an approach for reasoning about such discrepancies based on synthetic time-series scRNA-seq data generated by variational autoencoders. The synthetic dynamical patterns induced in a low-dimensional manifold reflect biologically plausible temporal patterns, such as dividing cell clusters during a differentiation process. We consider manifolds from different dimensionality reduction techniques, such as principal component analysis, t-distributed stochastic neighbor embedding, uniform manifold approximation, and projection and single-cell variational inference. We illustrate how the proposed approach allows for reasoning about to what extent low-dimensional manifolds, obtained from different techniques, can capture different dynamical patterns. None of these techniques was found to be consistently superior and the results indicate that they may not reliably represent dynamics when used in isolation, underscoring the need to compare multiple perspectives. Thus, the proposed synthetic dynamical pattern approach provides a foundation for guiding future methods development to detect complex patterns in time-series scRNA-seq data. Maren Hackenberg, Laia Canal Guitart, Rolf Backofen, Harald Binder |
Briefings Bioinform. | 3 |
| 2025 | An evolutionary approach to predict the orientation of CRISPR arraysabstractCRISPR-Cas is a defense system of bacteria and archaea against phages. Parts of the foreign DNA, called spacers, are incorporated into the CRISPR array which constitutes the immune memory. The orientation of CRISPR arrays is crucial for analyzing and understanding the functionality of CRISPR systems and their targets. Several methods have been developed to identify the orientation of a CRISPR array. To predict the orientation, different methods use different features such as the repeat sequences between the spacers, the location of the leader sequence, the Cas genes, or PAMs. However, those features are often not sufficient to predict the orientation with certainty, or different methods disagree. Remarkably, almost all CRISPR systems have been found to insert spacers in a polarized manner at the leader end of the array. We introduce CRISPR-evOr, a method that leverages the resulting patterns to predict the acquisition orientation for (a group of) CRISPR arrays by reconstructing and comparing the likelihood of their evolutionary history with respect to both possible acquisition orientations. The new method is independent of Cas type, leader existence and location, and transcription orientation. CRISPR-evOr is thus particularly useful for arrays that other CRISPR orientation tools cannot predict confidently and to verify or resolve conflicting predictions from existing tools. CRISPR-evOr currently confidently predicts the orientation of 28.3% of the arrays in the considered subset of CRISPRCasdb, which other tools like CRISPRDirection and CRISPRstrand cannot reliably orient. As our tool leverages evolutionary information we expect this percentage to grow in the future when more closely related arrays will be available. Additionally, CRISPR-evOr provides confident decisions for rare subtypes of CRISPR arrays, where knowledge about repeats and leaders and their orientation is limited. Axel Fehrenbach, Alexander Mitrofanov, Omer S. Alkhnbashi, Rolf Backofen, Franz Baumdicker |
PLoS Comput. Biol. | 4 |
| 2024 | CRISPert: A Transformer-Based Model for CRISPR-Cas Off-Target Prediction
William Jobson Pargeter, Rolf Backofen |
ECML/PKDD (7) | 2 |
| 2024 | Partial RNA designabstractMOTIVATION: RNA design is a key technique to achieve new functionality in fields like synthetic biology or biotechnology. Computational tools could help to find such RNA sequences but they are often limited in their formulation of the search space. RESULTS: In this work, we propose partial RNA design, a novel RNA design paradigm that addresses the limitations of current RNA design formulations. Partial RNA design describes the problem of designing RNAs from arbitrary RNA sequences and structure motifs with multiple design goals. By separating the design space from the objectives, our formulation enables the design of RNAs with variable lengths and desired properties, while still allowing precise control over sequence and structure constraints at individual positions. Based on this formulation, we introduce a new algorithm, libLEARNA, capable of efficiently solving different constraint RNA design tasks. A comprehensive analysis of various problems, including a realistic riboswitch design task, reveals the outstanding performance of libLEARNA and its robustness. AVAILABILITY AND IMPLEMENTATION: libLEARNA is open-source and publicly available at: https://github.com/automl/learna_tools. Frederic Runge, Jörg K. H. Franke, Daniel Fertmann, Rolf Backofen, Frank Hutter |
Bioinform. | 4 |
| 2023 | RNA interaction format: a general data format for RNA interactionsabstractSUMMARY: RNA molecules play crucial roles in various biological processes. They mediate their function mainly by interacting with other RNAs or proteins. At present, information about these interactions is distributed over different resources, often providing the data in simple tab-delimited formats that differ between the databases. There is no standardized data format that can capture the nature of all these different interactions in detail. AVAILABILITY AND IMPLEMENTATION: Here, we propose the RNA interaction format (RIF) for the detailed representation of RNA-RNA and RNA-Protein interactions and provide reference implementations in C/C++, Python, and JavaScript. RIF is released under licence GNU General Public License version 3 (GNU GPLv3) and is available on https://github.com/RNABioInfo/rna-interaction-format. Richard A. Schäfer, Dominik Rabsch, Guillaume E. Scholz, Peter F. Stadler, Wolfgang R. Hess, Rolf Backofen, Jörg Fallmann, Björn Voß |
Bioinform. | 6 |
| 2023 | Transformer-based tool recommendation system in GalaxyabstractBACKGROUND: Galaxy is a web-based open-source platform for scientific analyses. Researchers use thousands of high-quality tools and workflows for their respective analyses in Galaxy. Tool recommender system predicts a collection of tools that can be used to extend an analysis. In this work, a tool recommender system is developed by training a transformer on workflows available on Galaxy Europe and its performance is compared to other neural networks such as recurrent, convolutional and dense neural networks. RESULTS: The transformer neural network achieves two times faster convergence, has significantly lower model usage (model reconstruction and prediction) time and shows a better generalisation that goes beyond training workflows than the older tool recommender system created using RNN in Galaxy. In addition, the transformer also outperforms CNN and DNN on several key indicators. It achieves a faster convergence time, lower model usage time, and higher quality tool recommendations than CNN. Compared to DNN, it converges faster to a higher precision@k metric (approximately 0.98 by transformer compared to approximately 0.9 by DNN) and shows higher quality tool recommendations. CONCLUSION: Our work shows a novel usage of transformers to recommend tools for extending scientific workflows. A more robust tool recommendation model, created using a transformer, having significantly lower usage time than RNN and CNN, higher precision@k than DNN, and higher quality tool recommendations than all three neural networks, will benefit researchers in creating scientifically significant workflows and exploratory data analysis in Galaxy. Additionally, the ability to train faster than all three neural networks imparts more scalability for training on larger datasets consisting of millions of tool sequences. Open-source scripts to create the recommendation model are available under MIT licence at https://github.com/anuprulez/galaxy_tool_recommendation_transformers. Björn A. Grüning, Rolf Backofen |
BMC Bioinform. | 3 |
| 2022 | RiboReport - benchmarking tools for ribosome profiling-based identification of open reading frames in bacteriaabstractSmall proteins encoded by short open reading frames (ORFs) with 50 codons or fewer are emerging as an important class of cellular macromolecules in diverse organisms. However, they often evade detection by proteomics or in silico methods. Ribosome profiling (Ribo-seq) has revealed widespread translation in genomic regions previously thought to be non-coding, driving the development of ORF detection tools using Ribo-seq data. However, only a handful of tools have been designed for bacteria, and these have not yet been systematically compared. Here, we aimed to identify tools that use Ribo-seq data to correctly determine the translational status of annotated bacterial ORFs and also discover novel translated regions with high sensitivity. To this end, we generated a large set of annotated ORFs from four diverse bacterial organisms, manually labeled for their translation status based on Ribo-seq data, which are available for future benchmarking studies. This set was used to investigate the predictive performance of seven Ribo-seq-based ORF detection tools (REPARATION_blast, DeepRibo, Ribo-TISH, PRICE, smORFer, ribotricer and SPECtre), as well as IRSOM, which uses coding potential and RNA-seq coverage only. DeepRibo and REPARATION_blast robustly predicted translated ORFs, including sORFs, with no significant difference for ORFs in close proximity to other genes versus stand-alone genes. However, no tool predicted a set of novel, experimentally verified sORFs with high sensitivity. Start codon predictions with smORFer show the value of initiation site profiling data to further improve the sensitivity of ORF prediction tools in bacteria. Overall, we find that bacterial tools perform well for sORF detection, although there is potential for improving their performance, applicability, usability and reproducibility. Rick Gelhausen, Teresa Müller, Sarah L. Svensson, Omer S. Alkhnbashi, Cynthia M. Sharma, Florian Eggenhofer, Rolf Backofen |
Briefings Bioinform. | 7 |
| 2022 | CRISPRtracrRNA: robust approach for CRISPR tracrRNA detectionabstractMOTIVATION: The CRISPR-Cas9 system is a Type II CRISPR system that has rapidly become the most versatile and widespread tool for genome engineering. It consists of two components, the Cas9 effector protein, and a single guide RNA that combines the spacer (for identifying the target) with the tracrRNA, a trans-activating small RNA required for both crRNA maturation and interference. While there are well-established methods for screening Cas effector proteins and CRISPR arrays, the detection of tracrRNA remains the bottleneck in detecting Class 2 CRISPR systems. RESULTS: We introduce a new pipeline CRISPRtracrRNA for screening and evaluation of tracrRNA candidates in genomes. This pipeline combines evidence from different components of the Cas9-sgRNA complex. The core is a newly developed structural model via covariance models from a sequence-structure alignment of experimentally validated tracrRNAs. As additional evidence, we determine the terminator signal (required for the tracrRNA transcription) and the RNA-RNA interaction between the CRISPR array repeat and the 5'-part of the tracrRNA. Repeats are detected via an ML-based approach (CRISPRidenify). Providing further evidence, we detect the cassette containing the Cas9 (Type II CRISPR systems) and Cas12 (Type V CRISPR systems) effector protein. Our tool is the first for detecting tracrRNA for Type V systems. AVAILABILITY AND IMPLEMENTATION: The implementation of the CRISPRtracrRNA is available on GitHub upon requesting the access permission, (https://github.com/BackofenLab/CRISPRtracrRNA). Data generated in this study can be obtained upon request to the corresponding person: Rolf Backofen ([email protected]). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Mitrofanov, Marcus Ziemann, Omer S. Alkhnbashi, Wolfgang R. Hess, Rolf Backofen |
Bioinform. | 5 |
| 2022 | Peakhood: individual site context extraction for CLIP-seq peak regionsabstractMOTIVATION: CLIP-seq is by far the most widely used method to determine transcriptome-wide binding sites of RNA-binding proteins (RBPs). The binding site locations are identified from CLIP-seq read data by tools termed peak callers. Many RBPs bind to a spliced RNA (i.e. transcript) context, but all currently available peak callers only consider and report the genomic context. To accurately model protein binding behavior, a tool is needed for the individual context assignment to CLIP-seq peak regions. RESULTS: Here we present Peakhood, the first tool that utilizes CLIP-seq peak regions identified by peak callers, in tandem with CLIP-seq read information and genomic annotations, to determine which context applies, individually for each peak region. For sites assigned to transcript context, it further determines the most likely splice variant, and merges results for any number of datasets to obtain a comprehensive collection of transcript context binding sites. AVAILABILITY AND IMPLEMENTATION: Peakhood is freely available under MIT license at: https://github.com/BackofenLab/Peakhood. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael Uhl, Dominik Rabsch, Florian Eggenhofer, Rolf Backofen |
Bioinform. | 4 |
| 2021 | HRIBO: high-throughput analysis of bacterial ribosome profiling dataabstractMOTIVATION: Ribosome profiling (Ribo-seq) is a powerful approach based on deep sequencing of cDNA libraries generated from ribosome-protected RNA fragments to explore the translatome of a cell, and is especially useful for the detection of small proteins (50-100 amino acids) that are recalcitrant to many standard biochemical and in silico approaches. While pipelines are available to analyze Ribo-seq data, none are designed explicitly for the automatic processing and analysis of data from bacteria, nor are they focused on the discovery of unannotated open reading frames (ORFs). RESULTS: We present HRIBO (High-throughput annotation by Ribo-seq), a workflow to enable reproducible and high-throughput analysis of bacterial Ribo-seq data. The workflow performs all required pre-processing and quality control steps. Importantly, HRIBO outputs annotation-independent ORF predictions based on two complementary bacteria-focused tools, and integrates them with additional feature information and expression values. This facilitates the rapid and high-confidence discovery of novel ORFs and their prioritization for functional characterization. AVAILABILITY AND IMPLEMENTATION: HRIBO is a free and open source project available under the GPL-3 license at: https://github.com/RickGelhausen/HRIBO. Rick Gelhausen, Sarah L. Svensson, Kathrin Froschauer, Florian Heyl, Lydia Hadjeras, Cynthia M. Sharma, Florian Eggenhofer, Rolf Backofen |
Bioinform. | 8 |
| 2021 | pyGenomeTracks: reproducible plots for multivariate genomic datasetsabstractMOTIVATION: Generating publication ready plots to display multiple genomic tracks can pose a serious challenge. Making desirable and accurate figures requires considerable effort. This is usually done by hand or using a vector graphic software. RESULTS: pyGenomeTracks (PGT) is a modular plotting tool that easily combines multiple tracks. It enables a reproducible and standardized generation of highly customizable and publication ready images. AVAILABILITY AND IMPLEMENTATION: PGT is available through a graphical interface on https://usegalaxy.eu and through the command line. It is provided on conda via the bioconda channel, on pip and it is openly developed on github: https://github.com/deeptools/pyGenomeTracks. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lucille Lopez-Delisle, Leily Rabbani, Joachim Wolff, Vivek Bhardwaj 0002, Rolf Backofen, Björn A. Grüning, Fidel Ramírez, Thomas Manke |
Bioinform. | 5 |
| 2021 | Casboundary: automated definition of integral Cas cassettesabstractMOTIVATION: CRISPR-Cas are important systems found in most archaeal and many bacterial genomes, providing adaptive immunity against mobile genetic elements in prokaryotes. The CRISPR-Cas systems are encoded by a set of consecutive cas genes, here termed cassette. The identification of cassette boundaries is key for finding cassettes in CRISPR research field. This is often carried out by using Hidden Markov Models and manual annotation. In this article, we propose the first method able to automatically define the cassette boundaries. In addition, we present a Cas-type predictive model used by the method to assign each gene located in the region defined by a cassette's boundaries a Cas label from a set of pre-defined Cas types. Furthermore, the proposed method can detect potentially new cas genes and decompose a cassette into its modules. RESULTS: We evaluate the predictive performance of our proposed method on data collected from the two most recent CRISPR classification studies. In our experiments, we obtain an average similarity of 0.86 between the predicted and expected cassettes. Besides, we achieve F-scores above 0.9 for the classification of cas genes of known types and 0.73 for the unknown ones. Finally, we conduct two additional study cases, where we investigate the occurrence of potentially new cas genes and the occurrence of module exchange between different genomes. AVAILABILITY AND IMPLEMENTATION: https://github.com/BackofenLab/Casboundary. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Victor Alexandre Padilha, Omer S. Alkhnbashi, Shiraz A. Shah, André C. P. L. F. de Carvalho, Rolf Backofen |
Bioinform. | 6 |
| 2021 | A SARS-CoV-2 sequence submission tool for the European Nucleotide ArchiveabstractSUMMARY: Many aspects of the global response to the COVID-19 pandemic are enabled by the fast and open publication of SARS-CoV-2 genetic sequence data. The European Nucleotide Archive (ENA) is the European recommended open repository for genetic sequences. In this work, we present a tool for submitting raw sequencing reads of SARS-CoV-2 to ENA. The tool features a single-step submission process, a graphical user interface, tabular-formatted metadata and the possibility to remove human reads prior to submission. A Galaxy wrap of the tool allows users with little or no bioinformatics knowledge to do bulk sequencing read submissions. The tool is also packed in a Docker container to ease deployment. AVAILABILITY AND IMPLEMENTATION: CLI ENA upload tool is available at github.com/usegalaxy-eu/ena-upload-cli (DOI 10.5281/zenodo.4537621); Galaxy ENA upload tool at toolshed.g2.bx.psu.edu/view/iuc/ena_upload/382518f24d6d and github.com/galaxyproject/tools-iuc/tree/master/tools/ena_upload (development); and ENA upload Galaxy container at github.com/ELIXIR-Belgium/ena-upload-container (DOI 10.5281/zenodo.4730785). Miguel Roncoroni, Bert Droesbeke, Ignacio Eguinoa, Kim De Ruyck, Flora D'anna, Dilmurat Yusuf, Björn A. Grüning, Rolf Backofen, Frederik Coppens |
Bioinform. | 8 |
| 2021 | Scool: a new data storage format for single-cell Hi-C dataabstractBioinformatics (2020) doi: 10.1093/bioinformatics/btaa924 In the originally published version of this manuscript, requested author amendments to Figure 1. were inadvertently omitted prior to publishing. This error has now been corrected. Joachim Wolff, Nezar Abdennur, Rolf Backofen, Björn A. Grüning |
Bioinform. | 3 |
| 2021 | Scool: a new data storage format for single-cell Hi-C dataabstractMOTIVATION: Single-cell Hi-C research currently lacks an efficient, easy to use and shareable data storage format. Recent studies have used a variety of sub-optimal solutions: publishing raw data only, text-based interaction matrices, or reusing established Hi-C storage formats for single interaction matrices. These approaches are storage and pre-processing intensive, require long labour time and are often error-prone. RESULTS: The single-cell cooler file format (scool) provides an efficient, user-friendly and storage-saving approach for single-cell Hi-C data. It is a flavour of the established cooler format and guarantees stable API support. AVAILABILITY AND IMPLEMENTATION: The single-cell cooler format is part of the cooler file format as of API version 0.8.9. It is available via pip, conda and github: https://github.com/mirnylab/cooler. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joachim Wolff, Nezar Abdennur, Rolf Backofen, Björn A. Grüning |
Bioinform. | 3 |
| 2021 | Robust and efficient single-cell Hi-C clustering with approximate k-nearest neighbor graphsabstractMOTIVATION: Hi-C technology provides insights into the 3D organization of the chromatin, and the single-cell Hi-C method enables researchers to gain knowledge about the chromatin state in individual cell levels. Single-cell Hi-C interaction matrices are high dimensional and very sparse. To cluster thousands of single-cell Hi-C interaction matrices, they are flattened and compiled into one matrix. Depending on the resolution, this matrix can have a few million or even billions of features; therefore, computations can be memory intensive. We present a single-cell Hi-C clustering approach using an approximate nearest neighbors method based on locality-sensitive hashing to reduce the dimensions and the computational resources. RESULTS: The presented method can process a 10 kb single-cell Hi-C dataset with 2600 cells and needs 40 GB of memory, while competitive approaches are not computable even with 1 TB of memory. It can be shown that the differentiation of the cells by their chromatin folding properties and, therefore, the quality of the clustering of single-cell Hi-C data is advantageous compared to competitive algorithms. AVAILABILITY AND IMPLEMENTATION: The presented clustering algorithm is part of the scHiCExplorer, is available on Github https://github.com/joachimwolff/scHiCExplorer, and as a conda package via the bioconda channel. The approximate nearest neighbors implementation is available via https://github.com/joachimwolff/sparse-neighbors-search and as a conda package via the bioconda channel. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joachim Wolff, Rolf Backofen, Björn A. Grüning |
Bioinform. | 2 |
| 2020 | The locality dilemma of Sankoff-like RNA alignmentsabstractMOTIVATION: Elucidating the functions of non-coding RNAs by homology has been strongly limited due to fundamental computational and modeling issues. While existing simultaneous alignment and folding (SA&F) algorithms successfully align homologous RNAs with precisely known boundaries (global SA&F), the more pressing problem of identifying new classes of homologous RNAs in the genome (local SA&F) is intrinsically more difficult and much less understood. Typically, the length of local alignments is strongly overestimated and alignment boundaries are dramatically mispredicted. We hypothesize that local SA&F approaches are compromised this way due to a score bias, which is caused by the contribution of RNA structure similarity to their overall alignment score. RESULTS: In the light of this hypothesis, we study pairwise local SA&F for the first time systematically-based on a novel local RNA alignment benchmark set and quality measure. First, we vary the relative influence of structure similarity compared to sequence similarity. Putting more emphasis on the structure component leads to overestimating the length of local alignments. This clearly shows the bias of current scores and strongly hints at the structure component as its origin. Second, we study the interplay of several important scoring parameters by learning parameters for local and global SA&F. The divergence of these optimized parameter sets underlines the fundamental obstacles for local SA&F. Third, by introducing a position-wise correction term in local SA&F, we constructively solve its principal issues. AVAILABILITY AND IMPLEMENTATION: The benchmark data, detailed results and scripts are available at https://github.com/BackofenLab/local_alignment. The RNA alignment tool LocARNA, including the modifications proposed in this work, is available at https://github.com/s-will/LocARNA/releases/tag/v2.0.0RC6. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Teresa Müller, Milad Miladi, Frank Hutter, Ivo L. Hofacker, Sebastian Will, Rolf Backofen |
Bioinform. | 6 |
| 2020 | Heterogeneous networks integration for disease-gene prioritization with node kernelsabstractMOTIVATION: The identification of disease-gene associations is a task of fundamental importance in human health research. A typical approach consists in first encoding large gene/protein relational datasets as networks due to the natural and intuitive property of graphs for representing objects' relationships and then utilizing graph-based techniques to prioritize genes for successive low-throughput validation assays. Since different types of interactions between genes yield distinct gene networks, there is the need to integrate different heterogeneous sources to improve the reliability of prioritization systems. RESULTS: We propose an approach based on three phases: first, we merge all sources in a single network, then we partition the integrated network according to edge density introducing a notion of edge type to distinguish the parts and finally, we employ a novel node kernel suitable for graphs with typed edges. We show how the node kernel can generate a large number of discriminative features that can be efficiently processed by linear regularized machine learning classifiers. We report state-of-the-art results on 12 disease-gene associations and on a time-stamped benchmark containing 42 newly discovered associations. AVAILABILITY AND IMPLEMENTATION: Source code: https://github.com/dinhinfotech/DiGI.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alessandro Sperduti, Rolf Backofen, Fabrizio Costa |
Bioinform. | 3 |
| 2020 | The impact of various seed, accessibility and interaction constraints on sRNA target prediction- a systematic assessmentabstractBACKGROUND: Seed and accessibility constraints are core features to enable highly accurate sRNA target screens based on RNA-RNA interaction prediction. Currently, available tools provide different (sets of) constraints and default parameter sets. Thus, it is hard to impossible for users to estimate the influence of individual restrictions on the prediction results. RESULTS: Here, we present a systematic assessment of the impact of established and new constraints on sRNA target prediction both on a qualitative as well as computational level. This is done exemplarily based on the performance of IntaRNA, one of the most exact sRNA target prediction tools. IntaRNA provides various ways to constrain considered seed interactions, e.g. based on seed length, its accessibility, minimal unpaired probabilities, or energy thresholds, beside analogous constraints for the overall interaction. Thus, our results reveal the impact of individual constraints and their combinations. CONCLUSIONS: This provides both a guide for users what is important and recommendations for existing and upcoming sRNA target prediction approaches.We show on a large sRNA target screen benchmark data set that only by altering the parameter set, IntaRNA recovers 30% more verified interactions while becoming 5-times faster. This exemplifies the potential of seed, accessibility and interaction constraints for sRNA target prediction. Martin Raden, Teresa Müller, Stefan Mautner, Rick Gelhausen, Rolf Backofen |
BMC Bioinform. | 5 |
| 2019 | Progress Towards Graph Optimization: Efficient Learning of Vector to Graph Space Mappings
Stefan Mautner, Rolf Backofen, Fabrizio Costa |
ESANN | 2 |
| 2019 | Fast and Accurate Structure Probability Estimation for Simultaneous Alignment and Folding of RNAsabstractMotivation: Simultaneous alignment and folding (SA&F) of RNAs is the indispensable gold standard for inferring the structure of non-coding RNAs and their general analysis. The original algorithm, proposed by Sankoff, solves the theoretical problem exactly with a complexity of O(n^6) in the full energy model. Over the last two decades, several variants and improvements of the Sankoff algorithm have been proposed to reduce its extreme complexity by proposing simplified energy models or imposing restrictions on the predicted alignments. Results: Here we introduce a novel variant of Sankoff’s algorithm that reconciles the simplifications of PMcomp, namely moving from the full energy model to a simpler base pair-based model, with the accuracy of the loop-based full energy model. Instead of estimating pseudo-energies from unconditional base pair probabilities, our model calculates energies from conditional base pair probabilities that allow to accurately capture structure probabilities, which obey a conditional dependency. Supporting modifications with surgical precision, this model gives rise to the fast and highly accurate novel algorithm Pankov (Probabilistic Sankoff-like simultaneous alignment and folding of RNAs inspired by Markov chains). Pankov benefits from the speed-up of excluding unreliable base-pairing without compromising the loop-based free energy model of the Sankoff’s algorithm. We show that Pankov outperforms its predecessors LocARNA and SPARSE in folding quality and is faster than LocARNA. Pankov is developed as a branch of the LocARNA package and available at https://github.com/mmiladi/Pankov. Milad Miladi, Martin Raden, Sebastian Will, Rolf Backofen |
WABI | 4 |
| 2019 | ShaKer: RNA SHAPE prediction using graph kernelabstractSUMMARY: SHAPE experiments are used to probe the structure of RNA molecules. We present ShaKer to predict SHAPE data for RNA using a graph-kernel-based machine learning approach that is trained on experimental SHAPE information. While other available methods require a manually curated reference structure, ShaKer predicts reactivity data based on sequence input only and by sampling the ensemble of possible structures. Thus, ShaKer is well placed to enable experiment-driven, transcriptome-wide SHAPE data prediction to enable the study of RNA structuredness and to improve RNA structure and RNA-RNA interaction prediction. For performance evaluation, we use accuracy and accessibility comparing to experimental SHAPE data and competing methods. We can show that Shaker outperforms its competitors and is able to predict high quality SHAPE annotations even when no reference structure is provided. AVAILABILITY AND IMPLEMENTATION: ShaKer is freely available at https://github.com/BackofenLab/ShaKer. Stefan Mautner, Soheila Montaseri, Milad Miladi, Martin Raden, Fabrizio Costa, Rolf Backofen |
Bioinform. | 6 |
| 2019 | Integration of accessibility data from structure probing into RNA-RNA interaction predictionabstractSUMMARY: Experimental structure probing data has been shown to improve thermodynamics-based RNA secondary structure prediction. To this end, chemical reactivity information (as provided e.g. by SHAPE) is incorporated, which encodes whether or not individual nucleotides are involved in intra-molecular structure. Since inter-molecular RNA-RNA interactions are often confined to unpaired RNA regions, SHAPE data is even more promising to improve interaction prediction. Here, we show how such experimental data can be incorporated seamlessly into accessibility-based RNA-RNA interaction prediction approaches, as implemented in IntaRNA. This is possible via the computation and use of unpaired probabilities that incorporate the structure probing information. We show that experimental SHAPE data can significantly improve RNA-RNA interaction prediction. We evaluate our approach by investigating interactions of a spliceosomal U1 snRNA transcript with its target splice sites. When SHAPE data is incorporated, known target sites are predicted with increased precision and specificity. AVAILABILITY AND IMPLEMENTATION: https://github.com/BackofenLab/IntaRNA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Milad Miladi, Soheila Montaseri, Rolf Backofen, Martin Raden |
Bioinform. | 3 |
| 2018 | CMV: visualization for RNA and protein family models and their comparisonsabstractSummary: A standard method for the identification of novel RNAs or proteins is homology search via probabilistic models. One approach relies on the definition of families, which can be encoded as covariance models (CMs) or Hidden Markov Models (HMMs). While being powerful tools, their complexity makes it tedious to investigate them in their (default) tabulated form. This specifically applies to the interpretation of comparisons between multiple models as in family clans. The Covariance model visualization tools (CMV) visualize CMs or HMMs to: I) Obtain an easily interpretable representation of HMMs and CMs; II) Put them in context with the structural sequence alignments they have been created from; III) Investigate results of model comparisons and highlight regions of interest. Availability and implementation: Source code (http://www.github.com/eggzilla/cmv), web-service (http://rna.informatik.uni-freiburg.de/CMVS). Supplementary information: Supplementary data are available at Bioinformatics online. Florian Eggenhofer, Ivo L. Hofacker, Rolf Backofen, Christian Höner zu Siederdissen |
Bioinform. | 3 |
| 2018 | MechRNA: prediction of lncRNA mechanisms from RNA-RNA and RNA-protein interactionsabstractMotivation: Long non-coding RNAs (lncRNAs) are defined as transcripts longer than 200 nt that do not get translated into proteins. Often these transcripts are processed (spliced, capped and polyadenylated) and some are known to have important biological functions. However, most lncRNAs have unknown or poorly understood functions. Nevertheless, because of their potential role in cancer, lncRNAs are receiving a lot of attention, and the need for computational tools to predict their possible mechanisms of action is more than ever. Fundamentally, most of the known lncRNA mechanisms involve RNA-RNA and/or RNA-protein interactions. Through accurate predictions of each kind of interaction and integration of these predictions, it is possible to elucidate potential mechanisms for a given lncRNA. Results: Here, we introduce MechRNA, a pipeline for corroborating RNA-RNA interaction prediction and protein binding prediction for identifying possible lncRNA mechanisms involving specific targets or on a transcriptome-wide scale. The first stage uses a version of IntaRNA2 with added functionality for efficient prediction of RNA-RNA interactions with very long input sequences, allowing for large-scale analysis of lncRNA interactions with little or no loss of optimality. The second stage integrates protein binding information pre-computed by GraphProt, for both the lncRNA and the target. The final stage involves inferring the most likely mechanism for each lncRNA/target pair. This is achieved by generating candidate mechanisms from the predicted interactions, the relative locations of these interactions and correlation data, followed by selection of the most likely mechanistic explanation using a combined P-value. We applied MechRNA on a number of recently identified cancer-related lncRNAs (PCAT1, PCAT29 and ARLnc1) and also on two well-studied lncRNAs (PCA3 and 7SL). This led to the identification of hundreds of high confidence potential targets for each lncRNA and corresponding mechanisms. These predictions include the known competitive mechanism of 7SL with HuR for binding on the tumor suppressor TP53, as well as mechanisms expanding what is known about PCAT1 and ARLn1 and their targets BRCA2 and AR, respectively. For PCAT1-BRCA2, the mechanism involves competitive binding with HuR, which we confirmed using HuR immunoprecipitation assays. Availability and implementation: MechRNA is available for download at https://bitbucket.org/compbio/mechrna. Supplementary information: Supplementary data are available at Bioinformatics online. Alexander Gawronski, Michael Uhl, Yajia Zhang, Yen-Yi Lin, Yashar S. Niknafs, Varune R. Ramnarine, Rohit Malik, Felix Feng, Arul M. Chinnaiyan, Colin C. Collins, Süleyman Cenk Sahinalp, Rolf Backofen |
Bioinform. | 12 |
| 2017 | RNAscClust: clustering RNA sequences using structure conservation and graph based motifsabstractMOTIVATION: Clustering RNA sequences with common secondary structure is an essential step towards studying RNA function. Whereas structural RNA alignment strategies typically identify common structure for orthologous structured RNAs, clustering seeks to group paralogous RNAs based on structural similarities. However, existing approaches for clustering paralogous RNAs, do not take the compensatory base pair changes obtained from structure conservation in orthologous sequences into account. RESULTS: Here, we present RNAscClust , the implementation of a new algorithm to cluster a set of structured RNAs taking their respective structural conservation into account. For a set of multiple structural alignments of RNA sequences, each containing a paralog sequence included in a structural alignment of its orthologs, RNAscClust computes minimum free-energy structures for each sequence using conserved base pairs as prior information for the folding. The paralogs are then clustered using a graph kernel-based strategy, which identifies common structural features. We show that the clustering accuracy clearly benefits from an increasing degree of compensatory base pair changes in the alignments. AVAILABILITY AND IMPLEMENTATION: RNAscClust is available at http://www.bioinf.uni-freiburg.de/Software/RNAscClust . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Milad Miladi, Alexander Junge, Fabrizio Costa, Stefan E. Seemann, Jakob Hull Havgaard, Jan Gorodkin, Rolf Backofen |
Bioinform. | 7 |
| 2017 | Jupyter and Galaxy: Easing entry barriers into complex data analyses for biomedical researchersabstractWhat does it take to convert a heap of sequencing data into a publishable result? First, common tools are employed to reduce primary data (sequencing reads) to a form suitable for further analyses (i.e., the list of variable sites). The subsequent exploratory stage is much more ad hoc and requires the development of custom scripts and pipelines, making it problematic for biomedical researchers. Here, we describe a hybrid platform combining common analysis pathways with the ability to explore data interactively. It aims to fully encompass and simplify the "raw data-to-publication" pathway and make it reproducible. Björn A. Grüning, Eric Rasche, Boris Rebolledo-Jaramillo, Carl Eberhard, Torsten Houwaart, John Chilton, Nathan Coraor, Rolf Backofen, James Taylor 0001, Anton Nekrutenko |
PLoS Comput. Biol. | 8 |
| 2016 | AptaTRACE: Elucidating Sequence-Structure Binding Motifs by Uncovering Selection Trends in HT-SELEX Experiments
Phuong Dao, Jan Hoinka, Yijie Wang 0004, Mayumi Takahashi, Jiehua Zhou, Fabrizio Costa, John Rossi, John Burnett, Rolf Backofen, Teresa M. Przytycka |
RECOMB | 9 |
| 2016 | Characterizing leader sequences of CRISPR lociabstractMOTIVATION: The CRISPR-Cas system is an adaptive immune system in many archaea and bacteria, which provides resistance against invading genetic elements. The first phase of CRISPR-Cas immunity is called adaptation, in which small DNA fragments are excised from genetic elements and are inserted into a CRISPR array generally adjacent to its so called leader sequence at one end of the array. It has been shown that transcription initiation and adaptation signals of the CRISPR array are located within the leader. However, apart from promoters, there is very little knowledge of sequence or structural motifs or their possible functions. Leader properties have mainly been characterized through transcriptional initiation data from single organisms but large-scale characterization of leaders has remained challenging due to their low level of sequence conservation. RESULTS: We developed a method to successfully detect leader sequences by focusing on the consensus repeat of the adjacent CRISPR array and weak upstream conservation signals. We applied our tool to the analysis of a comprehensive genomic database and identified several characteristic properties of leader sequences specific to archaea and bacteria, ranging from distinctive sizes to preferential indel localization. CRISPRleader provides a full annotation of the CRISPR array, its strand orientation as well as conserved core leader boundaries that can be uploaded to any genome browser. In addition, it outputs reader-friendly HTML pages for conserved leader clusters from our database. AVAILABILITY AND IMPLEMENTATION: CRISPRleader and multiple sequence alignments for all 195 leader clusters are available at http://www.bioinf.uni-freiburg.de/Software/CRISPRleader/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Omer S. Alkhnbashi, Shiraz A. Shah, Roger A. Garrett, Sita J. Saunders, Fabrizio Costa, Rolf Backofen |
Bioinform. | 6 |
| 2016 | SnoReport 2.0: new features and a refined Support Vector Machine to improve snoRNA identificationabstractBACKGROUND: snoReport uses RNA secondary structure prediction combined with machine learning as the basis to identify the two main classes of small nucleolar RNAs, the box H/ACA snoRNAs and the box C/D snoRNAs. Here, we present snoReport 2.0, which substantially improves and extends in the original method by: extracting new features for both box C/D and H/ACA box snoRNAs; developing a more sophisticated technique in the SVM training phase with recent data from vertebrate organisms and a careful choice of the SVM parameters C and γ; and using updated versions of tools and databases used for the construction of the original version of snoReport. To validate the new version and to demonstrate its improved performance, we tested snoReport 2.0 in different organisms. RESULTS: Results of the training and test phases of boxes H/ACA and C/D snoRNAs, in both versions of snoReport, are discussed. Validation on real data was performed to evaluate the predictions of snoReport 2.0. Our program was applied to a set of previously annotated sequences, some of them experimentally confirmed, of humans, nematodes, drosophilids, platypus, chickens and leishmania. We significantly improved the predictions for vertebrates, since the training phase used information of these organisms, but H/ACA box snoRNAs identification was improved for the other ones. CONCLUSION: We presented snoReport 2.0, to predict H/ACA box and C/D box snoRNAs, an efficient method to find true positives and avoid false positives in vertebrate organisms. H/ACA box snoRNA classifier showed an F-score of 93 % (an improvement of 10 % regarding the previous version), while C/D box snoRNA classifier, an F-Score of 94 % (improvement of 14 %). Besides, both classifiers exhibited performance measures above 90 %. These results show that snoReport 2.0 avoid false positives and false negatives, allowing to predict snoRNAs with high quality. In the validation phase, snoReport 2.0 predicted 67.43 % of vertebrate organisms for both classes. For Nematodes and Drosophilids, 69 % and 76.67 %, for H/ACA box snoRNAs were predicted, respectively, showing that snoReport 2.0 is good to identify snoRNAs in vertebrates and also H/ACA box snoRNAs in invertebrates organisms. João Victor de Araújo Oliveira, Fabrizio Costa, Rolf Backofen, Peter F. Stadler, Maria Emília M. T. Walter, Jana Schor |
BMC Bioinform. | 3 |
| 2015 | antaRNA: ant colony-based RNA sequence designabstractMOTIVATION: RNA sequence design is studied at least as long as the classical folding problem. Although for the latter the functional fold of an RNA molecule is to be found ,: inverse folding tries to identify RNA sequences that fold into a function-specific target structure. In combination with RNA-based biotechnology and synthetic biology ,: reliable RNA sequence design becomes a crucial step to generate novel biochemical components. RESULTS: In this article ,: the computational tool antaRNA is presented. It is capable of compiling RNA sequences for a given structure that comply in addition with an adjustable full range objective GC-content distribution ,: specific sequence constraints and additional fuzzy structure constraints. antaRNA applies ant colony optimization meta-heuristics and its superior performance is shown on a biological datasets. AVAILABILITY AND IMPLEMENTATION: http://www.bioinf.uni-freiburg.de/Software/antaRNA CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Robert Kleinkauf, Martin Raden, Rolf Backofen |
Bioinform. | 3 |
| 2015 | SPARSE: quadratic time simultaneous alignment and folding of RNAs without sequence-based heuristicsabstractMOTIVATION: RNA-Seq experiments have revealed a multitude of novel ncRNAs. The gold standard for their analysis based on simultaneous alignment and folding suffers from extreme time complexity of [Formula: see text]. Subsequently, numerous faster 'Sankoff-style' approaches have been suggested. Commonly, the performance of such methods relies on sequence-based heuristics that restrict the search space to optimal or near-optimal sequence alignments; however, the accuracy of sequence-based methods breaks down for RNAs with sequence identities below 60%. Alignment approaches like LocARNA that do not require sequence-based heuristics, have been limited to high complexity ([Formula: see text] quartic time). RESULTS: Breaking this barrier, we introduce the novel Sankoff-style algorithm 'sparsified prediction and alignment of RNAs based on their structure ensembles (SPARSE)', which runs in quadratic time without sequence-based heuristics. To achieve this low complexity, on par with sequence alignment algorithms, SPARSE features strong sparsification based on structural properties of the RNA ensembles. Following PMcomp, SPARSE gains further speed-up from lightweight energy computation. Although all existing lightweight Sankoff-style methods restrict Sankoff's original model by disallowing loop deletions and insertions, SPARSE transfers the Sankoff algorithm to the lightweight energy model completely for the first time. Compared with LocARNA, SPARSE achieves similar alignment and better folding quality in significantly less time (speedup: 3.7). At similar run-time, it aligns low sequence identity instances substantially more accurate than RAF, which uses sequence-based heuristics. Sebastian Will, Christina Otto, Milad Miladi, Mathias Möhl, Rolf Backofen |
Bioinform. | 5 |
| 2015 | antaRNA - Multi-objective inverse folding of pseudoknot RNA using ant-colony optimizationabstractBACKGROUND: Many functional RNA molecules fold into pseudoknot structures, which are often essential for the formation of an RNA's 3D structure. Currently the design of RNA molecules, which fold into a specific structure (known as RNA inverse folding) within biotechnological applications, is lacking the feature of incorporating pseudoknot structures into the design. Hairpin-(H)- and kissing hairpin-(K)-type pseudoknots cover a wide range of biologically functional pseudoknots and can be represented on a secondary structure level. RESULTS: The RNA inverse folding program antaRNA, which takes secondary structure, target GC-content and sequence constraints as input, is extended to provide solutions for such H- and K-type pseudoknotted secondary structure constraint. We demonstrate the easy and flexible interchangeability of modules within the antaRNA framework by incorporating pKiss as structure prediction tool capable of predicting the mentioned pseudoknot types. The performance of the approach is demonstrated on a subset of the Pseudobase ++ dataset. CONCLUSIONS: This new service is available via a standalone version and is also part of the Freiburg RNA Tools webservice. Furthermore, antaRNA is available in Galaxy and is part of the RNA-workbench Docker image. Robert Kleinkauf, Torsten Houwaart, Rolf Backofen, Martin Raden |
BMC Bioinform. | 3 |
| 2014 | CRISPRstrand: predicting repeat orientations to determine the crRNA-encoding strand at CRISPR lociabstractMOTIVATION: The discovery of CRISPR-Cas systems almost 20 years ago rapidly changed our perception of the bacterial and archaeal immune systems. CRISPR loci consist of several repetitive DNA sequences called repeats, inter-spaced by stretches of variable length sequences called spacers. This CRISPR array is transcribed and processed into multiple mature RNA species (crRNAs). A single crRNA is integrated into an interference complex, together with CRISPR-associated (Cas) proteins, to bind and degrade invading nucleic acids. Although existing bioinformatics tools can recognize CRISPR loci by their characteristic repeat-spacer architecture, they generally output CRISPR arrays of ambiguous orientation and thus do not determine the strand from which crRNAs are processed. Knowledge of the correct orientation is crucial for many tasks, including the classification of CRISPR conservation, the detection of leader regions, the identification of target sites (protospacers) on invading genetic elements and the characterization of protospacer-adjacent motifs. RESULTS: We present a fast and accurate tool to determine the crRNA-encoding strand at CRISPR loci by predicting the correct orientation of repeats based on an advanced machine learning approach. Both the repeat sequence and mutation information were encoded and processed by an efficient graph kernel to learn higher-order correlations. The model was trained and tested on curated data comprising >4500 CRISPRs and yielded a remarkable performance of 0.95 AUC ROC (area under the curve of the receiver operator characteristic). In addition, we show that accurate orientation information greatly improved detection of conserved repeat sequence families and structure motifs. We integrated CRISPRstrand predictions into our CRISPRmap web server of CRISPR conservation and updated the latter to version 2.0. AVAILABILITY: CRISPRmap and CRISPRstrand are available at http://rna.informatik.uni-freiburg.de/CRISPRmap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Omer S. Alkhnbashi, Fabrizio Costa, Shiraz A. Shah, Roger A. Garrett, Sita J. Saunders, Rolf Backofen |
Bioinform. | 6 |
| 2014 | MoDPepInt: an interactive web server for prediction of modular domain-peptide interactionsabstractUNLABELLED: MoDPepInt (Modular Domain Peptide Interaction) is a new easy-to-use web server for the prediction of binding partners for modular protein domains. Currently, we offer models for SH2, SH3 and PDZ domains via the tools SH2PepInt, SH3PepInt and PDZPepInt, respectively. More specifically, our server offers predictions for 51 SH2 human domains and 69 SH3 human domains via single domain models, and predictions for 226 PDZ domains across several species, via 43 multidomain models. All models are based on support vector machines with different kernel functions ranging from polynomial, to Gaussian, to advanced graph kernels. In this way, we model non-linear interactions between amino acid residues. Results were validated on manually curated datasets achieving competitive performance against various state-of-the-art approaches. AVAILABILITY AND IMPLEMENTATION: The MoDPepInt server is available under the URL http://modpepint.informatik.uni-freiburg.de/. Kousik Kundu, Martin Raden, Fabrizio Costa, Rolf Backofen |
Bioinform. | 4 |
| 2014 | BlockClust: efficient clustering and classification of non-coding RNAs from short read RNA-seq profilesabstractSUMMARY: Non-coding RNAs (ncRNAs) play a vital role in many cellular processes such as RNA splicing, translation, gene regulation. However the vast majority of ncRNAs still have no functional annotation. One prominent approach for putative function assignment is clustering of transcripts according to sequence and secondary structure. However sequence information is changed by post-transcriptional modifications, and secondary structure is only a proxy for the true 3D conformation of the RNA polymer. A different type of information that does not suffer from these issues and that can be used for the detection of RNA classes, is the pattern of processing and its traces in small RNA-seq reads data. Here we introduce BlockClust, an efficient approach to detect transcripts with similar processing patterns. We propose a novel way to encode expression profiles in compact discrete structures, which can then be processed using fast graph-kernel techniques. We perform both unsupervised clustering and develop family specific discriminative models; finally we show how the proposed approach is scalable, accurate and robust across different organisms, tissues and cell lines. AVAILABILITY: The whole BlockClust galaxy workflow including all tool dependencies is available at http://toolshed.g2.bx.psu.edu/view/rnateam/blockclust_workflow. Pavankumar Videm, Dominic Rose, Fabrizio Costa, Rolf Backofen |
Bioinform. | 4 |
| 2014 | ExpaRNA-P: simultaneous exact pattern matching and folding of RNAsabstractBACKGROUND: Identifying sequence-structure motifs common to two RNAs can speed up the comparison of structural RNAs substantially. The core algorithm of the existent approach ExpaRNA solves this problem for a priori known input structures. However, such structures are rarely known; moreover, predicting them computationally is no rescue, since single sequence structure prediction is highly unreliable. RESULTS: The novel algorithm ExpaRNA-P computes exactly matching sequence-structure motifs in entire Boltzmann-distributed structure ensembles of two RNAs; thereby we match and fold RNAs simultaneously, analogous to the well-known "simultaneous alignment and folding" of RNAs. While this implies much higher flexibility compared to ExpaRNA, ExpaRNA-P has the same very low complexity (quadratic in time and space), which is enabled by its novel structure ensemble-based sparsification. Furthermore, we devise a generalized chaining algorithm to compute compatible subsets of ExpaRNA-P's sequence-structure motifs. Resulting in the very fast RNA alignment approach ExpLoc-P, we utilize the best chain as anchor constraints for the sequence-structure alignment tool LocARNA. ExpLoc-P is benchmarked in several variants and versus state-of-the-art approaches. In particular, we formally introduce and evaluate strict and relaxed variants of the problem; the latter makes the approach sensitive to compensatory mutations. Across a benchmark set of typical non-coding RNAs, ExpLoc-P has similar accuracy to LocARNA but is four times faster (in both variants), while it achieves a speed-up over 30-fold for the longest benchmark sequences (≈400nt). Finally, different ExpLoc-P variants enable tailoring of the method to specific application scenarios. ExpaRNA-P and ExpLoc-P are distributed as part of the LocARNA package. The source code is freely available at http://www.bioinf.uni-freiburg.de/Software/ExpaRNA-P . CONCLUSIONS: ExpaRNA-P's novel ensemble-based sparsification reduces its complexity to quadratic time and space. Thereby, ExpaRNA-P significantly speeds up sequence-structure alignment while maintaining the alignment quality. Different ExpaRNA-P variants support a wide range of applications. Christina Otto, Mathias Möhl, Steffen Heyne, Mika Amit, Gad M. Landau, Rolf Backofen, Sebastian Will |
BMC Bioinform. | 6 |
| 2014 | Local Exact Pattern Matching for Non-FixedRNA StructuresabstractDetecting local common sequence-structure regions of RNAs is a biologically important problem. Detecting such regions allows biologists to identify functionally relevant similarities between the inspected molecules. We developed dynamic programming algorithms for finding common structure-sequence patterns between two RNAs. The RNAs are given by their sequence and a set of potential base pairs with associated probabilities. In contrast to prior work on local pattern matching of RNAs, we support the breaking of arcs. This allows us to add flexibility over matching only fixed structures; potentially matching only a similar subset of specified base pairs. We present an O(n(3)) algorithm for local exact pattern matching between two nested RNAs, and an O(n(3) log n) algorithm for one nested RNA and one bounded-unlimited RNA. In addition, an algorithm for approximate pattern matching is introduced that for two given nested RNAs and a number k, finds the maximal local pattern matching score between the two RNAs with at most k mismatches in O(n(3)k(2)) time. Finally, we present an O(n(3)) algorithm for finding the most similar subforest between two nested RNAs. Mika Amit, Rolf Backofen, Steffen Heyne, Gad M. Landau, Mathias Möhl, Christina Otto, Sebastian Will |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | Atom Mapping with Constraint Programming
Martin Raden, Feras Nahar, Heinz Ekker, Rolf Backofen, Peter F. Stadler, Christoph Flamm |
CP | 4 |
| 2013 | SPARSE: Quadratic Time Simultaneous Alignment and Folding of RNAs without Sequence-Based Heuristics
Sebastian Will, Christina Schmiedl, Milad Miladi, Mathias Möhl, Rolf Backofen |
RECOMB | 5 |
| 2013 | Distribution of Graph-Distances in Boltzmann Ensembles of RNA Secondary Structures
Rolf Backofen, Markus Fricke, Manja Marz, Jing Qin 0006, Peter F. Stadler |
WABI | 1 |
| 2013 | A graph kernel approach for alignment-free domain-peptide interaction prediction with an application to human SH3 domainsabstractMOTIVATION: State-of-the-art experimental data for determining binding specificities of peptide recognition modules (PRMs) is obtained by high-throughput approaches like peptide arrays. Most prediction tools applicable to this kind of data are based on an initial multiple alignment of the peptide ligands. Building an initial alignment can be error-prone, especially in the case of the proline-rich peptides bound by the SH3 domains. RESULTS: Here, we present a machine-learning approach based on an efficient graph-kernel technique to predict the specificity of a large set of 70 human SH3 domains, which are an important class of PRMs. The graph-kernel strategy allows us to (i) integrate several types of physico-chemical information for each amino acid, (ii) consider high-order correlations between these features and (iii) eliminate the need for an initial peptide alignment. We build specialized models for each human SH3 domain and achieve competitive predictive performance of 0.73 area under precision-recall curve, compared with 0.27 area under precision-recall curve for state-of-the-art methods based on position weight matrices. We show that better models can be obtained when we use information on the noninteracting peptides (negative examples), which is currently not used by the state-of-the art approaches based on position weight matrices. To this end, we analyze two strategies to identify subsets of high confidence negative data. The techniques introduced here are more general and hence can also be used for any other protein domains, which interact with short peptides (i.e. other PRMs). AVAILABILITY: The program with the predictive models can be found at http://www.bioinf.uni-freiburg.de/Software/SH3PepInt/SH3PepInt.tar.gz. We also provide a genome-wide prediction for all 70 human SH3 domains, which can be found under http://www.bioinf.uni-freiburg.de/Software/SH3PepInt/Genome-Wide-Predictions.tar.gz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kousik Kundu, Fabrizio Costa, Rolf Backofen |
Bioinform. | 3 |
| 2012 | Local Exact Pattern Matching for Non-fixed RNA Structures
Mika Amit, Rolf Backofen, Steffen Heyne, Gad M. Landau, Mathias Möhl, Christina Schmiedl, Sebastian Will |
CPM | 2 |
| 2012 | Impact of the Energy Model on the Complexity of RNA Folding with Pseudoknots
Saad Sheikh, Rolf Backofen, Yann Ponty |
CPM | 2 |
| 2012 | Exact Pattern Matching for RNA Structure Ensembles
Christina Schmiedl, Mathias Möhl, Steffen Heyne, Mika Amit, Gad M. Landau, Sebastian Will, Rolf Backofen |
RECOMB | 7 |
| 2012 | GraphClust: alignment-free structural clustering of local RNA secondary structuresabstractMOTIVATION: Clustering according to sequence-structure similarity has now become a generally accepted scheme for ncRNA annotation. Its application to complete genomic sequences as well as whole transcriptomes is therefore desirable but hindered by extremely high computational costs. RESULTS: We present a novel linear-time, alignment-free method for comparing and clustering RNAs according to sequence and structure. The approach scales to datasets of hundreds of thousands of sequences. The quality of the retrieved clusters has been benchmarked against known ncRNA datasets and is comparable to state-of-the-art sequence-structure methods although achieving speedups of several orders of magnitude. A selection of applications aiming at the detection of novel structural ncRNAs are presented. Exemplarily, we predicted local structural elements specific to lincRNAs likely functionally associating involved transcripts to vital processes of the human nervous system. In total, we predicted 349 local structural RNA elements. AVAILABILITY: The GraphClust pipeline is available on request. Steffen Heyne, Fabrizio Costa, Dominic Rose, Rolf Backofen |
Bioinform. | 4 |
| 2012 | Navigating the unexplored seascape of pre-miRNA candidates in single-genome approachesabstractMOTIVATION: The computational search for novel microRNA (miRNA) precursors often involves some sort of structural analysis with the aim of identifying which type of structures are prone to being recognized and processed by the cellular miRNA-maturation machinery. A natural way to tackle this problem is to perform clustering over the candidate structures along with known miRNA precursor structures. Mixed clusters allow then the identification of candidates that are similar to known precursors. Given the large number of pre-miRNA candidates that can be identified in single-genome approaches, even after applying several filters for precursor robustness and stability, a conventional structural clustering approach is unfeasible. RESULTS: We propose a method to represent candidate structures in a feature space, which summarizes key sequence/structure characteristics of each candidate. We demonstrate that proximity in this feature space is related to sequence/structure similarity, and we select candidates that have a high similarity to known precursors. Additional filtering steps are then applied to further reduce the number of candidates to those with greater transcriptional potential. Our method is compared with another single-genome method (TripletSVM) in two datasets, showing better performance in one and comparable performance in the other, for larger training sets. Additionally, we show that our approach allows for a better interpretation of the results. AVAILABILITY AND IMPLEMENTATION: The MinDist method is implemented using Perl scripts and is freely available at http://www.cravela.org/?mindist=1. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nuno D. Mendes, Steffen Heyne, Ana T. Freitas, Marie-France Sagot, Rolf Backofen |
Bioinform. | 5 |
| 2011 | Computational discovery of human coding and non-coding transcripts with conserved splice sitesabstractMOTIVATION: Long non-coding RNAs (lncRNAs) resemble protein-coding mRNAs but do not encode proteins. Most lncRNAs are under lower sequence constraints than protein-coding genes and lack conserved secondary structures, making it hard to predict them computationally. RESULTS: We introduce an approach to predict spliced lncRNAs in vertebrate genomes combining comparative genomics and machine learning. It is based on detecting signatures of characteristic splice site evolution in vertebrate whole genome alignments. First, we predict individual splice sites, then assemble compatible sites into exon candidates, and finally predict multi-exon transcripts. Using a novel method to evaluate typical splice site substitution patterns that explicitly takes the species phylogeny into account, we show that individual splice sites can be accurately predicted. Since our approach relies only on predicted splice sites, it can uncover both coding and non-coding exons. We show that our predicted exons and partial transcripts are mostly non-coding and lack conserved secondary structures. These exons are of particular interest, since existing computational approaches cannot detect them. Transcriptome sequencing data indicate tissue-specific expression patterns of predicted exons and there is evidence that increasing sequencing depth and breadth will validate additional predictions. We also found a significant enrichment of predicted exons that form multi-exon transcript parts, and we experimentally validate such a novel multi-exon gene. Overall, we obtain 336 novel multi-exon transcript predictions from human intergenic regions. Our results indicate the existence of novel human transcripts that are conserved in evolution and our approach contributes to the completion of the human transcript catalog. AVAILABILITY AND IMPLEMENTATION: Predicted human splice sites, exons and gene structures together with a Perl implementation of the tree-based log-odds scoring and a supplementary PDF file containing additional figures and tables are available at: http://www.bioinf.uni-leipzig.de/publications/supplements/10-010. The five experimentally confirmed partial transcript isoforms have been deposited in GenBank under accession numbers HM587422-HM587426. Dominic Rose, Michael Hiller, Katharina Schutt, Jörg Hackermüller, Rolf Backofen, Peter F. Stadler |
Bioinform. | 5 |
| 2011 | PETcofold: predicting conserved interactions and structures of two multiple alignments of RNA sequencesabstractMOTIVATION: Predicting RNA-RNA interactions is essential for determining the function of putative non-coding RNAs. Existing methods for the prediction of interactions are all based on single sequences. Since comparative methods have already been useful in RNA structure determination, we assume that conserved RNA-RNA interactions also imply conserved function. Of these, we further assume that a non-negligible amount of the existing RNA-RNA interactions have also acquired compensating base changes throughout evolution. We implement a method, PETcofold, that can take covariance information in intra-molecular and inter-molecular base pairs into account to predict interactions and secondary structures of two multiple alignments of RNA sequences. RESULTS: PETcofold's ability to predict RNA-RNA interactions was evaluated on a carefully curated dataset of 32 bacterial small RNAs and their targets, which was manually extracted from the literature. For evaluation of both RNA-RNA interaction and structure prediction, we were able to extract only a few high-quality examples: one vertebrate small nucleolar RNA and four bacterial small RNAs. For these we show that the prediction can be improved by our comparative approach. Furthermore, PETcofold was evaluated on controlled data with phylogenetically simulated sequences enriched for covariance patterns at the interaction sites. We observed increased performance with increased amounts of covariance. AVAILABILITY: The program PETcofold is available as source code and can be downloaded from http://rth.dk/resources/petcofold. Stefan E. Seemann, Andreas S. Richter, Tanja Gesell, Rolf Backofen, Jan Gorodkin |
Bioinform. | 4 |
| 2011 | Structator: fast index-based search for RNA sequence-structure patternsabstractBACKGROUND: The secondary structure of RNA molecules is intimately related to their function and often more conserved than the sequence. Hence, the important task of searching databases for RNAs requires to match sequence-structure patterns. Unfortunately, current tools for this task have, in the best case, a running time that is only linear in the size of sequence databases. Furthermore, established index data structures for fast sequence matching, like suffix trees or arrays, cannot benefit from the complementarity constraints introduced by the secondary structure of RNAs. RESULTS: We present a novel method and readily applicable software for time efficient matching of RNA sequence-structure patterns in sequence databases. Our approach is based on affix arrays, a recently introduced index data structure, preprocessed from the target database. Affix arrays support bidirectional pattern search, which is required for efficiently handling the structural constraints of the pattern. Structural patterns like stem-loops can be matched inside out, such that the loop region is matched first and then the pairing bases on the boundaries are matched consecutively. This allows to exploit base pairing information for search space reduction and leads to an expected running time that is sublinear in the size of the sequence database. The incorporation of a new chaining approach in the search of RNA sequence-structure patterns enables the description of molecules folding into complex secondary structures with multiple ordered patterns. The chaining approach removes spurious matches from the set of intermediate results, in particular of patterns with little specificity. In benchmark experiments on the Rfam database, our method runs up to two orders of magnitude faster than previous methods. CONCLUSIONS: The presented method's sublinear expected running time makes it well suited for RNA sequence-structure pattern matching in large sequence databases. RNA molecules containing several stem-loop substructures can be described by multiple sequence-structure patterns and their matches are efficiently handled by a novel chaining method. Beyond our algorithmic contributions, we provide with Structator a complete and robust open-source software solution for index-based search of RNA sequence-structure patterns. The Structator software is available at http://www.zbh.uni-hamburg.de/Structator. Fernando Meyer, Stefan Kurtz, Rolf Backofen, Sebastian Will, Michael Beckstette |
BMC Bioinform. | 3 |
| 2010 | Time and Space Efficient RNA-RNA Interaction Prediction via Sparse Folding
Raheleh Salari, Mathias Möhl, Sebastian Will, Süleyman Cenk Sahinalp, Rolf Backofen |
RECOMB | 5 |
| 2010 | Sparsification of RNA Structure Prediction Including Pseudoknots
Mathias Möhl, Raheleh Salari, Sebastian Will, Rolf Backofen, Süleyman Cenk Sahinalp |
WABI | 4 |
| 2010 | Seed-based INTARNA prediction combined with GFP-reporter system identifies mRNA targets of the small RNA Yfr1abstractMOTIVATION: Prochlorococcus possesses the smallest genome of all sequenced photoautotrophs. Although the number of regulatory proteins in the genome is very small, the relative number of small regulatory RNAs is comparable with that of other bacteria. The compact genome size of Prochlorococcus offers an ideal system to search for targets of small RNAs (sRNAs) and to refine existing target prediction algorithms. RESULTS: Target predictions for the cyanobacterial sRNA Yfr1 were carried out with INTARNA in Prochlorococcus MED4. The ultraconserved Yfr1 sequence motif was defined as the putative interaction seed. To study the impact of Yfr1 on its predicted mRNA targets, a reporter system based on green fluorescent protein (GFP) was applied. We show that Yfr1 inhibits the translation of two predicted targets. We used mutation analysis to confirm that Yfr1 directly regulates its targets by an antisense interaction sequestering the ribosome binding site, and to assess the importance of interaction site accessibility. Andreas S. Richter, Christian Schleberger, Rolf Backofen, Claudia Steglich |
Bioinform. | 3 |
| 2010 | TassDB2 - A comprehensive database of subtle alternative splicing eventsabstractBACKGROUND: Subtle alternative splicing events involving tandem splice sites separated by a short (2-12 nucleotides) distance are frequent and evolutionarily widespread in eukaryotes, and a major contributor to the complexity of transcriptomes and proteomes. However, these events have been either omitted altogether in databases on alternative splicing, or only the cases of experimentally confirmed alternative splicing have been reported. Thus, a database which covers all confirmed cases of subtle alternative splicing as well as the numerous putative tandem splice sites (which might be confirmed once more transcript data becomes available), and allows to search for tandem splice sites with specific features and download the results, is a valuable resource for targeted experimental studies and large-scale bioinformatics analyses of tandem splice sites. Towards this goal we recently set up TassDB (Tandem Splice Site DataBase, version 1), which stores data about alternative splicing events at tandem splice sites separated by 3 nt in eight species. DESCRIPTION: We have substantially revised and extended TassDB. The currently available version 2 contains extensive information about tandem splice sites separated by 2-12 nt for the human and mouse transcriptomes including data on the conservation of the tandem motifs in five vertebrates. TassDB2 offers a user-friendly interface to search for specific genes or for genes containing tandem splice sites with specific features as well as the possibility to download result datasets. For example, users can search for cases of alternative splicing where the proportion of EST/mRNA evidence supporting the minor isoform exceeds a specific threshold, or where the difference in splice site scores is specified by the user. The predicted impact of each event on the protein is also reported, along with information about being a putative target for the nonsense-mediated decay (NMD) pathway. Links are provided to the UCSC genome browser and other external resources. CONCLUSION: TassDB2, available via http://www.tassdb.info, provides comprehensive resources for researchers interested in both targeted experimental studies and large-scale bioinformatics analyses of short distance tandem splice sites. Rileen Sinha, Thorsten Lenser, Niels Jahn, Ulrike Gausmann, Swetlana Friedel, Karol Szafranski, Klaus Huse, Philip Rosenstiel, Jochen Hampe, Stefan Schuster, Michael Hiller, Rolf Backofen, Matthias Platzer |
BMC Bioinform. | 12 |
| 2009 | Fast RNA Structure Alignment for Crossing Input Structures
Rolf Backofen, Gad M. Landau, Mathias Möhl, Dekel Tsur, Oren Weimann |
CPM | 1 |
| 2009 | Sparse RNA Folding: Time and Space Efficient Algorithms
Rolf Backofen, Dekel Tsur, Shay Zakov, Michal Ziv-Ukelson |
CPM | 1 |
| 2009 | Lifting Prediction to Alignment of RNA Pseudoknots
Mathias Möhl, Sebastian Will, Rolf Backofen |
RECOMB | 3 |
| 2009 | Simultaneous Alignment and Folding of Protein Sequences
Jérôme Waldispühl, Charles W. O'Donnell, Sebastian Will, Srini Devadas, Rolf Backofen, Bonnie Berger |
RECOMB | 5 |
| 2009 | biRNA: Fast RNA-RNA Binding Sites Prediction
Hamidreza Chitsaz, Rolf Backofen, Süleyman Cenk Sahinalp |
WABI | 2 |
| 2009 | Fast Prediction of RNA-RNA Interaction
Raheleh Salari, Rolf Backofen, Süleyman Cenk Sahinalp |
WABI | 2 |
| 2009 | A partition function algorithm for interacting nucleic acid strandsabstractUNLABELLED: Recent interests, such as RNA interference and antisense RNA regulation, strongly motivate the problem of predicting whether two nucleic acid strands interact. MOTIVATION: Regulatory non-coding RNAs (ncRNAs) such as microRNAs play an important role in gene regulation. Studies on both prokaryotic and eukaryotic cells show that such ncRNAs usually bind to their target mRNA to regulate the translation of corresponding genes. The specificity of these interactions depends on the stability of intermolecular and intramolecular base pairing. While methods like deep sequencing allow to discover an ever increasing set of ncRNAs, there are no high-throughput methods available to detect their associated targets. Hence, there is an increasing need for precise computational target prediction. In order to predict base-pairing probability of any two bases in interacting nucleic acids, it is necessary to compute the interaction partition function over the whole ensemble. The partition function is a scalar value from which various thermodynamic quantities can be derived. For example, the equilibrium concentration of each complex nucleic acid species and also the melting temperature of interacting nucleic acids can be calculated based on the partition function of the complex. RESULTS: We present a model for analyzing the thermodynamics of two interacting nucleic acid strands considering the most general type of interactions studied in the literature. We also present a corresponding dynamic programming algorithm that computes the partition function over (almost) all physically possible joint secondary structures formed by two interacting nucleic acids in O(n(6)) time. We verify the predictive power of our algorithm by computing (i) the melting temperature for interacting RNA pairs studied in the literature and (ii) the equilibrium concentration for several variants of the OxyS-fhlA complex. In both experiments, our algorithm shows high accuracy and outperforms competitors. AVAILABILITY: Software and web server is available at http://compbio.cs.sfu.ca/taverna/pirna/. SUPPLEMENTARY INFORMATION: Supplementary data are avaliable at Bioinformatics online. Hamidreza Chitsaz, Raheleh Salari, Süleyman Cenk Sahinalp, Rolf Backofen |
Bioinform. | 4 |
| 2009 | Lightweight comparison of RNAs based on exact sequence-structure matchesabstractMOTIVATION: Specific functions of ribonucleic acid (RNA) molecules are often associated with different motifs in the RNA structure. The key feature that forms such an RNA motif is the combination of sequence and structure properties. In this article, we introduce a new RNA sequence-structure comparison method which maintains exact matching substructures. Existing common substructures are treated as whole unit while variability is allowed between such structural motifs. Based on a fast detectable set of overlapping and crossing substructure matches for two nested RNA secondary structures, our method ExpaRNA (exact pattern of alignment of RNA) computes the longest collinear sequence of substructures common to two RNAs in O(H.nm) time and O(nm) space, where H << n.m for real RNA structures. Applied to different RNAs, our method correctly identifies sequence-structure similarities between two RNAs. RESULTS: We have compared ExpaRNA with two other alignment methods that work with given RNA structures, namely RNAforester and RNA_align. The results are in good agreement, but can be obtained in a fraction of running time, in particular for larger RNAs. We have also used ExpaRNA to speed up state-of-the-art Sankoff-style alignment tools like LocARNA, and observe a tradeoff between quality and speed. However, we get a speedup of 4.25 even in the highest quality setting, where the quality of the produced alignment is comparable to that of LocARNA alone. AVAILABILITY: The presented algorithm is implemented in the program ExpaRNA, which is available from our website (http://www.bioinf.uni-freiburg.de/Software). Steffen Heyne, Sebastian Will, Michael Beckstette, Rolf Backofen |
Bioinform. | 4 |
| 2009 | CPSP-web-tools: a server for 3D lattice protein studiesabstractUNLABELLED: Studies on proteins are often restricted to highly simplified models to face the immense computational complexity of the associated problems. Constraint-based protein structure prediction (CPSP) tools is a package of very fast algorithms for ab initio optimal structure prediction and related problems in 3D HP-models [cubic and face centered cubic (FCC)]. Here, we present CPSP-web-tools, an interactive online interface of these programs for their immediate use. They include the first method for the direct prediction of optimal energies and structures in 3D HP side-chain models. This newest extension of the CPSP approach is described here for the first time. AVAILABILITY AND IMPLEMENTATION: Free access at http://cpsp.informatik.uni-freiburg.de Martin Raden, Cameron Smith, Mohamad Rabbath, Marlien Edwards, Sebastian Will, Rolf Backofen |
Bioinform. | 6 |
| 2009 | Fast Feature Subset Selection in Biological Sequence AnalysisabstractBiological research produces a wealth of measured data. Neither it is easy for biologists to postulate hypotheses about the behavior or structure of the observed entity because the relevant properties measured are not seen in the ocean of measurements. Nor is it easy to design machine learning algorithms to classify or cluster the data items for the same reason. Algorithms for automatically selecting a highly predictive subset of the measured features can help to overcome these difficulties. We present an efficient feature selection strategy which can be applied to arbitrary feature selection problems. The core technique is a new method for estimating the quality of subsets from previously calculated qualities for smaller subsets by minimizing the mean standard error of estimated values with an approach common to support vector machines. This method can be integrated in many feature subset search algorithms. We have applied it with sequential search algorithms and have been able to reduce the number of quality calculations for finding accurate feature subsets by about 70%. We show these improvements by applying our approach to the problem of finding highly predictive feature subsets for transcription factor binding sites. Rainer Pudimat, Rolf Backofen, Ernst Günter Schukat-Talamazzini |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2008 | Fixed Parameter Tractable Alignment of RNA Structures Including Arbitrary Pseudoknots
Mathias Möhl, Sebastian Will, Rolf Backofen |
CPM | 3 |
| 2008 | IntaRNA: efficient prediction of bacterial sRNA targets incorporating target site accessibility and seed regionsabstractMOTIVATION: During the last few years, several new small regulatory RNAs (sRNAs) have been discovered in bacteria. Most of them act as post-transcriptional regulators by base pairing to a target mRNA, causing translational repression or activation, or mRNA degradation. Numerous sRNAs have already been identified, but the number of experimentally verified targets is considerably lower. Consequently, computational target prediction is in great demand. Many existing target prediction programs neglect the accessibility of target sites and the existence of a seed, while other approaches are either specialized to certain types of RNAs or too slow for genome-wide searches. RESULTS: We introduce INTARNA, a new general and fast approach to the prediction of RNA-RNA interactions incorporating accessibility of target sites as well as the existence of a user-definable seed. We successfully applied INTARNA to the prediction of bacterial sRNA targets and determined the exact locations of the interactions with a higher accuracy than competing programs. AVAILABILITY: http://www.bioinf.uni-freiburg.de/Software/ Anke Busch, Andreas S. Richter, Rolf Backofen |
Bioinform. | 3 |
| 2008 | CPSP-tools - Exact and complete algorithms for high-throughput 3D lattice protein studiesabstractBACKGROUND: The principles of protein folding and evolution pose problems of very high inherent complexity. Often these problems are tackled using simplified protein models, e.g. lattice proteins. The CPSP-tools package provides programs to solve exactly and completely the problems typical of studies using 3D lattice protein models. Among the tasks addressed are the prediction of (all) globally optimal and/or suboptimal structures as well as sequence design and neutral network exploration. RESULTS: In contrast to stochastic approaches, which are not capable of answering many fundamental questions, our methods are based on fast, non-heuristic techniques. The resulting tools are designed for high-throughput studies of 3D-lattice proteins utilising the Hydrophobic-Polar (HP) model. The source bundle is freely available 1. CONCLUSION: The CPSP-tools package is the first set of exact and complete methods for extensive, high-throughput studies of non-restricted 3D-lattice protein models. In particular, our package deals with cubic and face centered cubic (FCC) lattices. Martin Raden, Sebastian Will, Rolf Backofen |
BMC Bioinform. | 3 |
| 2008 | Improved identification of conserved cassette exons using Bayesian networksabstractBACKGROUND: Alternative splicing is a major contributor to the diversity of eukaryotic transcriptomes and proteomes. Currently, large scale detection of alternative splicing using expressed sequence tags (ESTs) or microarrays does not capture all alternative splicing events. Moreover, for many species genomic data is being produced at a far greater rate than corresponding transcript data, hence in silico methods of predicting alternative splicing have to be improved. RESULTS: Here, we show that the use of Bayesian networks (BNs) allows accurate prediction of evolutionary conserved exon skipping events. At a stringent false positive rate of 0.5%, our BN achieves an improved true positive rate of 61%, compared to a previously reported 50% on the same dataset using support vector machines (SVMs). Incorporating several novel discriminative features such as intronic splicing regulatory elements leads to the improvement. Features related to mRNA secondary structure increase the prediction performance, corroborating previous findings that secondary structures are important for exon recognition. Random labelling tests rule out overfitting. Cross-validation on another dataset confirms the increased performance. When using the same dataset and the same set of features, the BN matches the performance of an SVM in earlier literature. Remarkably, we could show that about half of the exons which are labelled constitutive but receive a high probability of being alternative by the BN, are in fact alternative exons according to the latest EST data. Finally, we predict exon skipping without using conservation-based features, and achieve a true positive rate of 29% at a false positive rate of 0.5%. CONCLUSION: BNs can be used to achieve accurate identification of alternative exons and provide clues about possible dependencies between relevant features. The near-identical performance of the BN and SVM when using the same features shows that good classification depends more on features than on the choice of classifier. Conservation based features continue to be the most informative, and hence distinguishing alternative exons from constitutive ones without using conservation based features remains a challenging problem. Rileen Sinha, Michael Hiller, Rainer Pudimat, Ulrike Gausmann, Matthias Platzer, Rolf Backofen |
BMC Bioinform. | 6 |
| 2007 | Inferring Noncoding RNA Families and Classes by Means of Genome-Scale Structure-Based ClusteringabstractThe RFAM database defines families of ncRNAs by means of sequence similarities that are sufficient to establish homology. In some cases, such as microRNAs and box H/ACA snoRNAs, functional commonalities define classes of RNAs that are characterized by structural similarities, and typically consist of multiple RNA families. Recent advances in high-throughput transcriptomics and comparative genomics have produced very large sets of putative noncoding RNAs and regulatory RNA signals. For many of them, evidence for stabilizing selection acting on their secondary structures has been derived, and at least approximate models of their structures have been computed. The overwhelming majority of these hypothetical RNAs cannot be assigned to established families or classes. We present here a structure-based clustering approach that is capable of extracting putative RNA classes from genome-wide surveys for structured RNAs. The LocARNA (local alignment of RNA) tool implements a novel variant of the Sankoff algorithm that is sufficiently fast to deal with several thousand candidate sequences. The method is also robust against false positive predictions, i.e., a contamination of the input data with unstructured or nonconserved sequences. We have successfully tested the LocARNA-based clustering approach on the sequences of the RFAM-seed alignments. Furthermore, we have applied it to a previously published set of 3,332 predicted structured elements in the Ciona intestinalis genome (Missal K, Rose D, Stadler PF (2005) Noncoding RNAs in Ciona intestinalis. Bioinformatics 21 (Supplement 2): i77-i78). In addition to recovering, e.g., tRNAs as a structure-based class, the method identifies several RNA families, including microRNA and snoRNA candidates, and suggests several novel classes of ncRNAs for which to date no representative has been experimentally characterized. Sebastian Will, Kristin Reiche, Ivo L. Hofacker, Peter F. Stadler, Rolf Backofen |
PLoS Comput. Biol. | 5 |
| 2006 | Local Alignment of RNA Sequences with Arbitrary Scoring Schemes
Rolf Backofen, Danny Hermelin, Gad M. Landau, Oren Weimann |
CPM | 1 |
| 2006 | INFO-RNA - a fast approach to inverse RNA foldingabstractMOTIVATION: The structure of RNA molecules is often crucial for their function. Therefore, secondary structure prediction has gained much interest. Here, we consider the inverse RNA folding problem, which means designing RNA sequences that fold into a given structure. RESULTS: We introduce a new algorithm for the inverse folding problem (INFO-RNA) that consists of two parts; a dynamic programming method for good initial sequences and a following improved stochastic local search that uses an effective neighbor selection method. During the initialization, we design a sequence that among all sequences adopts the given structure with the lowest possible energy. For the selection of neighbors during the search, we use a kind of look-ahead of one selection step applying an additional energy-based criterion. Afterwards, the pre-ordered neighbors are tested using the actual optimization criterion of minimizing the structure distance between the target structure and the mfe structure of the considered neighbor. We compared our algorithm to RNAinverse and RNA-SSD for artificial and biological test sets. Using INFO-RNA, we performed better than RNAinverse and in most cases, we gained better results than RNA-SSD, the probably best inverse RNA folding tool on the market. AVAILABILITY: www.bioinf.uni-freiburg.de?Subpages/software.html. Anke Busch, Rolf Backofen |
Bioinform. | 2 |
| 2005 | A New Distance Measure of RNA Ensembles and Its Application to Phylogenetic Tree Construction
Sven Siebert, Rolf Backofen |
CIBCB | 2 |
| 2005 | Normalized Similarity of RNA Sequences
Rolf Backofen, Danny Hermelin, Gad M. Landau, Oren Weimann |
SPIRE | 1 |
| 2005 | SECISDesign: a server to design SECIS-elements within the coding sequenceabstractSUMMARY: SECISDesign is a server for the design of SECIS-elements and arbitrary RNA-elements within the coding sequence of an mRNA. The element has to satisfy both structure and sequence constraints. At the same time, a certain amino acid similarity to the original protein has to be kept. The designed sequence can be used for recombinant expression of selenoproteins in Escherichia coli. AVAILABILITY: The server is available at http://www.bio.inf.uni-jena.de/Software/SECISDesign/index.html. Anke Busch, Sebastian Will, Rolf Backofen |
Bioinform. | 3 |
| 2005 | A multiple-feature framework for modelling and predicting transcription factor binding sitesabstractMOTIVATION: The identification of transcription factor binding sites in promoter sequences is an important problem, since it reveals information about the transcriptional regulation of genes. For analysing transcriptional regulation, computational approaches for predicting putative binding sites are applied. Commonly used stochastic models for binding sites are position-specific score matrices, which show weak predictive power. RESULTS: We have developed a probabilistic modelling approach, which allows to consider diverse characteristic binding site properties to obtain more accurate representations of binding sites. These properties are modelled as random variables in Bayesian networks, which are capable of dealing with dependencies among binding site properties. Cross-validation on several datasets shows improvements in the false positive error rate and the significance (P-value) of true binding sites. Rainer Pudimat, Ernst Günter Schukat-Talamazzini, Rolf Backofen |
Bioinform. | 3 |
| 2005 | MARNA: multiple alignment and consensus structure prediction of RNAs based on sequence structure comparisonsabstractMOTIVATION: Due to the importance of considering secondary structures in aligning functional RNAs, several pairwise sequence-structure alignment methods have been developed. They use extended alignment scores that evaluate secondary structure information in addition to sequence information. However, two problems for the multiple alignment step remain. First, how to combine pairwise sequence-structure alignments into a multiple alignment and second, how to generate secondary structure information for sequences whose explicit structural information is missing. RESULTS: We describe a novel approach for multiple alignment of RNAs (MARNA) taking into consideration both the primary and the secondary structures. It is based on pairwise sequence-structure comparisons of RNAs. From these sequence-structure alignments, libraries of weighted alignment edges are generated. The weights reflect the sequential and structural conservation. For sequences whose secondary structures are missing, the libraries are generated by sampling low energy conformations. The libraries are then processed by the T-Coffee system, which is a consistency based multiple alignment method. Furthermore, we are able to extract a consensus-sequence and -structure from a multiple alignment. We have successfully tested MARNA on several datasets taken from the Rfam database. Sven Siebert, Rolf Backofen |
Bioinform. | 2 |
| 2004 | Computational Design of New and Recombinant Selenoproteins
Rolf Backofen, Anke Busch |
CPM | 1 |
| 2004 | Fast Detection of Common Sequence Structure Patterns in RNAs
Rolf Backofen, Sven Siebert |
SPIRE | 1 |
| 2003 | A Constraint-Based Approach to Structure Prediction for Simplified Protein Models That Outperforms Other Existing Methods
Rolf Backofen, Sebastian Will |
ICLP | 1 |
| 2002 | On the Complexity of Protein Similarity Search under mRNA Structure Constraints
Rolf Backofen, N. S. Narayanaswamy, Firas Swidan |
STACS | 1 |
| 2001 | Fast, Constraint-Based Threading of HP-Sequences to Hydrophobic Cores
Rolf Backofen, Sebastian Will |
CP | 1 |
| 2001 | Optimally Compact Finite Sphere Packings - Hydrophobic Cores in the FCC
Rolf Backofen, Sebastian Will |
CPM | 1 |
| 2001 | RNA-Sequence-Structure Properties and Selenocysteine Insertion
Rolf Backofen |
IDA | 1 |
| 2000 | An Upper Bound for Number of Contacts in the HP-Model on the Face-Centered-Cubic Lattice (FCC)
Rolf Backofen |
CPM | 1 |
| 1999 | Excluding Symmetries in Constraint-Based Search
Rolf Backofen, Sebastian Will |
CP | 1 |
| 1999 | Application of constraint programming techniques for structure prediction of lattice proteins with extended alphabetsabstractMOTIVATION: Predicting the ground state of biopolymers is a notoriously hard problem in biocomputing. Model systems, such as lattice proteins, are simple tools and valuable to test and improve new methods. Best known are models with sequences composed from a binary (hydrophobic and polar) alphabet. The major drawback is the degeneracy, i.e. the number of different ground state conformations. RESULTS: We show how recently developed constraint programming techniques can be used to solve the structure prediction problem efficiently for a higher order alphabet. To our knowledge it is the first report of an exact and computationally feasible solution to model proteins of length up to 36 and without resorting to maximally compact states. We further show that degeneracy is reduced by more than one order of magnitude and that ground state conformations are not necessarily compact. Therefore, more realistic protein simulations become feasible with our model. Rolf Backofen, Sebastian Will, Erich Bornberg-Bauer |
Bioinform. | 1 |
| 1998 | Constraint Techniques for Solving the Protein Structure Prediction Problem
Rolf Backofen |
CP | 1 |
| 1998 | How to Win a Game with Features
Rolf Backofen, Ralf Treinen |
Inf. Comput. | 1 |
| 1996 | Controlling Functional Uncertainty
Rolf Backofen |
ECAI | 1 |
| 1995 | A Complete and Recursive Feature Theory
Rolf Backofen, Gert Smolka |
Theor. Comput. Sci. | 1 |
| 1994 | DISCO-An HPSG-based NLP System and its Application for Appointment Scheduling Project Note
Hans Uszkoreit, Rolf Backofen, Stephan Busemann, Abdel Kader Diagne, Elizabeth A. Hinkelman, Walter Kasper, Bernd Kiefer, Hans-Ulrich Krieger, Klaus Netter, Günter Neumann, Stephan Oepen, Stephen P. Spackman |
COLING | 2 |
| 1994 | Regular Path Expressions in Feature Logic
Rolf Backofen |
J. Symb. Comput. | 1 |
| 1993 | On the Decidability of Functional UncertaintyabstractWe show that feature logic extended by functional uncertainty is decidable, even if one admits cyclic descriptions. We present an algorithm, which solves feature descriptions containing functional uncertainty in two phases, both phases using a set of deterministic and non-deterministic rewrite rules. We then compare our algorithm with the one of Kaplan and Maxwell, that does not cover cyclic feature descriptions. Rolf Backofen |
ACL | 1 |
| 1993 | A Complete and Recursive Feature TheoryabstractVarious feature descriptions are being employed in constrained-based grammar formalisms. The common notational primitive of these descriptions are functional attributes called features. The descriptions considered in this paper are the possibly quantified first-order formulae obtained from a signature of features and sorts. We establish a complete first-order theory FT by means of three axiom schemes and construct three elementarily equivalent models.One of the models consists of so-called feature graphs, a data structure common in computational linguistics. The other two models consist of so-called feature trees, a record-like data structure generalizing the trees corresponding to first-order terms.Our completeness proof exhibits a terminating simplification system deciding validity and satisfiability of possibly quantified feature descriptions. Rolf Backofen, Gert Smolka |
ACL | 1 |
| 1993 | Regular Path Expression in Feature Logic
Rolf Backofen |
RTA | 1 |