VLDB 2026 Research / reviewers in the wild / expert
Ivo L. Hofacker
dblp:99/6107
· DBLP profile ↗
43ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0001-7132-0800ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Old Dog, New Tricks: Exact Seeding Strategy Improves RNA Design Performances
Théo Boury, Leonhard Sidl, Ivo L. Hofacker, Yann Ponty, Hua-Ting Yao |
RECOMB | 3 |
| 2023 | Phylogenetic Information as Soft Constraints in RNA Secondary Structure Prediction
Sarah von Löhneysen, Thomas Spicher, Yuliia Varenyk, Hua-Ting Yao, Ronny Lorenz, Ivo L. Hofacker, Peter F. Stadler |
ISBRA | 6 |
| 2023 | DrTransformer: heuristic cotranscriptional RNA folding using the nearest neighbor energy modelabstractMOTIVATION: Folding during transcription can have an important influence on the structure and function of RNA molecules, as regions closer to the 5' end can fold into metastable structures before potentially stronger interactions with the 3' end become available. Thermodynamic RNA folding models are not suitable to predict structures that result from cotranscriptional folding, as they can only calculate properties of the equilibrium distribution. Other software packages that simulate the kinetic process of RNA folding during transcription exist, but they are mostly applicable for short sequences. RESULTS: We present a new algorithm that tracks changes to the RNA secondary structure ensemble during transcription. At every transcription step, new representative local minima are identified, a neighborhood relation is defined and transition rates are estimated for kinetic simulations. After every simulation, a part of the ensemble is removed and the remainder is used to search for new representative structures. The presented algorithm is deterministic (up to numeric instabilities of simulations), fast (in comparison with existing methods), and it is capable of folding RNAs much longer than 200 nucleotides. AVAILABILITY AND IMPLEMENTATION: This software is open-source and available at https://github.com/ViennaRNA/drtransformer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Stefan Badelt, Ronny Lorenz, Ivo L. Hofacker |
Bioinform. | 3 |
| 2023 | Modified RNAs and predictions with the ViennaRNA PackageabstractMOTIVATION: In living organisms, many RNA molecules are modified post-transcriptionally. This turns the widely known four-letter RNA alphabet ACGU into a much larger one with currently more than 300 known distinct modified bases. The roles for the majority of modified bases remain uncertain, but many are already well-known for their ability to influence the preferred structures that an RNA may adopt. In fact, tRNAs sometimes require certain modifications to fold into their cloverleaf shaped structure. However, predicting the structure of RNAs with base modifications is still difficult due to the lack of efficient algorithms that can deal with the extended sequence alphabet, as well as missing parameter sets that account for the changes in stability induced by the modified bases. RESULTS: We present an approach to include sparse energy parameter data for modified bases into the ViennaRNA Package. Our method does not require any changes to the underlying efficient algorithms but instead uses a set of plug-in constraints that adapt the predictions in terms of loop evaluation at runtime. These adaptations are efficient in the sense that they are only performed for loops where additional parameters are actually available for. In addition, our approach also facilitates the inclusion of more modified bases as soon as further parameters become available. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are available at https://www.tbi.univie.ac.at/RNA. Yuliia Varenyk, Thomas Spicher, Ivo L. Hofacker, Ronny Lorenz |
Bioinform. | 3 |
| 2021 | RNAxplorer: harnessing the power of guiding potentials to sample RNA landscapesabstractMOTIVATION: Predicting the folding dynamics of RNAs is a computationally difficult problem, first and foremost due to the combinatorial explosion of alternative structures in the folding space. Abstractions are therefore needed to simplify downstream analyses, and thus make them computationally tractable. This can be achieved by various structure sampling algorithms. However, current sampling methods are still time consuming and frequently fail to represent key elements of the folding space. METHOD: We introduce RNAxplorer, a novel adaptive sampling method to efficiently explore the structure space of RNAs. RNAxplorer uses dynamic programming to perform an efficient Boltzmann sampling in the presence of guiding potentials, which are accumulated into pseudo-energy terms and reflect similarity to already well-sampled structures. This way, we effectively steer sampling toward underrepresented or unexplored regions of the structure space. RESULTS: We developed and applied different measures to benchmark our sampling methods against its competitors. Most of the measures show that RNAxplorer produces more diverse structure samples, yields rare conformations that may be inaccessible to other sampling methods and is better at finding the most relevant kinetic traps in the landscape. Thus, it produces a more representative coarse graining of the landscape, which is well suited to subsequently compute better approximations of RNA folding kinetics. AVAILABILITYAND IMPLEMENTATION: https://github.com/ViennaRNA/RNAxplorer/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gregor Entzian, Ivo L. Hofacker, Yann Ponty, Ronny Lorenz, Andrea Tanzer |
Bioinform. | 2 |
| 2020 | The locality dilemma of Sankoff-like RNA alignmentsabstractMOTIVATION: Elucidating the functions of non-coding RNAs by homology has been strongly limited due to fundamental computational and modeling issues. While existing simultaneous alignment and folding (SA&F) algorithms successfully align homologous RNAs with precisely known boundaries (global SA&F), the more pressing problem of identifying new classes of homologous RNAs in the genome (local SA&F) is intrinsically more difficult and much less understood. Typically, the length of local alignments is strongly overestimated and alignment boundaries are dramatically mispredicted. We hypothesize that local SA&F approaches are compromised this way due to a score bias, which is caused by the contribution of RNA structure similarity to their overall alignment score. RESULTS: In the light of this hypothesis, we study pairwise local SA&F for the first time systematically-based on a novel local RNA alignment benchmark set and quality measure. First, we vary the relative influence of structure similarity compared to sequence similarity. Putting more emphasis on the structure component leads to overestimating the length of local alignments. This clearly shows the bias of current scores and strongly hints at the structure component as its origin. Second, we study the interplay of several important scoring parameters by learning parameters for local and global SA&F. The divergence of these optimized parameter sets underlines the fundamental obstacles for local SA&F. Third, by introducing a position-wise correction term in local SA&F, we constructively solve its principal issues. AVAILABILITY AND IMPLEMENTATION: The benchmark data, detailed results and scripts are available at https://github.com/BackofenLab/local_alignment. The RNA alignment tool LocARNA, including the modifications proposed in this work, is available at https://github.com/s-will/LocARNA/releases/tag/v2.0.0RC6. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Teresa Müller, Milad Miladi, Frank Hutter, Ivo L. Hofacker, Sebastian Will, Rolf Backofen |
Bioinform. | 4 |
| 2018 | CMV: visualization for RNA and protein family models and their comparisonsabstractSummary: A standard method for the identification of novel RNAs or proteins is homology search via probabilistic models. One approach relies on the definition of families, which can be encoded as covariance models (CMs) or Hidden Markov Models (HMMs). While being powerful tools, their complexity makes it tedious to investigate them in their (default) tabulated form. This specifically applies to the interpretation of comparisons between multiple models as in family clans. The Covariance model visualization tools (CMV) visualize CMs or HMMs to: I) Obtain an easily interpretable representation of HMMs and CMs; II) Put them in context with the structural sequence alignments they have been created from; III) Investigate results of model comparisons and highlight regions of interest. Availability and implementation: Source code (http://www.github.com/eggzilla/cmv), web-service (http://rna.informatik.uni-freiburg.de/CMVS). Supplementary information: Supplementary data are available at Bioinformatics online. Florian Eggenhofer, Ivo L. Hofacker, Rolf Backofen, Christian Höner zu Siederdissen |
Bioinform. | 2 |
| 2017 | RNAblueprint: flexible multiple target nucleic acid sequence designabstractMOTIVATION: Realizing the value of synthetic biology in biotechnology and medicine requires the design of molecules with specialized functions. Due to its close structure to function relationship, and the availability of good structure prediction methods and energy models, RNA is perfectly suited to be synthetically engineered with predefined properties. However, currently available RNA design tools cannot be easily adapted to accommodate new design specifications. Furthermore, complicated sampling and optimization methods are often developed to suit a specific RNA design goal, adding to their inflexibility. RESULTS: We developed a C ++ library implementing a graph coloring approach to stochastically sample sequences compatible with structural and sequence constraints from the typically very large solution space. The approach allows to specify and explore the solution space in a well defined way. Our library also guarantees uniform sampling, which makes optimization runs performant by not only avoiding re-evaluation of already found solutions, but also by raising the probability of finding better solutions for long optimization runs. We show that our software can be combined with any other software package to allow diverse RNA design applications. Scripting interfaces allow the easy adaption of existing code to accommodate new scenarios, making the whole design process very flexible. We implemented example design approaches written in Python to demonstrate these advantages. AVAILABILITY AND IMPLEMENTATION: RNAblueprint , Python implementations and benchmark datasets are available at github: https://github.com/ViennaRNA . CONTACT: [email protected], [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Stefan Hammer, Birgit Tschiatschek, Christoph Flamm, Ivo L. Hofacker, Sven Findeiß |
Bioinform. | 4 |
| 2016 | Computational Design of a Circular RNA with Prionlike BehaviorabstractRNA molecules engineered to fold into predefined conformations have enabled the design of a multitude of functional RNA devices in the field of synthetic biology and nanotechnology. More complex designs require efficient computational methods, which need to consider not only equilibrium thermodynamics but also the kinetics of structure formation. Here we present a novel type of RNA design that mimics the behavior of prions, that is, sequences capable of interaction-triggered autocatalytic replication of conformations. Our design was computed with the ViennaRNA package and is based on circular RNA that embeds domains amenable to intermolecular kissing interactions. Stefan Badelt, Christoph Flamm, Ivo L. Hofacker |
Artif. Life | 3 |
| 2016 | Pseudoknots in RNA folding landscapesabstractMOTIVATION: The function of an RNA molecule is not only linked to its native structure, which is usually taken to be the ground state of its folding landscape, but also in many cases crucially depends on the details of the folding pathways such as stable folding intermediates or the timing of the folding process itself. To model and understand these processes, it is necessary to go beyond ground state structures. The study of rugged RNA folding landscapes holds the key to answer these questions. Efficient coarse-graining methods are required to reduce the intractably vast energy landscapes into condensed representations such as barrier trees or basin hopping graphs : BHG) that convey an approximate but comprehensive picture of the folding kinetics. So far, exact and heuristic coarse-graining methods have been mostly restricted to the pseudoknot-free secondary structures. Pseudoknots, which are common motifs and have been repeatedly hypothesized to play an important role in guiding folding trajectories, were usually excluded. RESULTS: We generalize the BHG framework to include pseudoknotted RNA structures and systematically study the differences in predicted folding behavior depending on whether pseudoknotted structures are allowed to occur as folding intermediates or not. We observe that RNAs with pseudoknotted ground state structures tend to have more pseudoknotted folding intermediates than RNAs with pseudoknot-free ground state structures. The occurrence and influence of pseudoknotted intermediates on the folding pathway, however, appear to depend very strongly on the individual RNAs so that no general rule can be inferred. AVAILABILITY AND IMPLEMENTATION: The algorithms described here are implemented in C++ as standalone programs. Its source code and Supplemental material can be freely downloaded from http://www.tbi.univie.ac.at/bhg.html. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marcel Kucharík, Ivo L. Hofacker, Peter F. Stadler, Jing Qin 0006 |
Bioinform. | 2 |
| 2016 | SHAPE directed RNA foldingabstractSUMMARY: Chemical mapping experiments allow for nucleotide resolution assessment of RNA structure. We demonstrate that different strategies of integrating probing data with thermodynamics-based RNA secondary structure prediction algorithms can be implemented by means of soft constraints. This amounts to incorporating suitable pseudo-energies into the standard energy model for RNA secondary structures. As a showcase application for this new feature of the ViennaRNA Package we compare three distinct, previously published strategies to utilize SHAPE reactivities for structure prediction. The new tool is benchmarked on a set of RNAs with known reference structure. AVAILABILITY AND IMPLEMENTATION: The capability for SHAPE directed RNA folding is part of the upcoming release of the ViennaRNA Package 2.2, for which a preliminary release is already freely available at http://www.tbi.univie.ac.at/RNA. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronny Lorenz, Dominik Luntzer, Ivo L. Hofacker, Peter F. Stadler, Michael T. Wolfinger |
Bioinform. | 3 |
| 2016 | Practical Guidelines for Incorporating Knowledge-Based and Data-Driven Strategies into the Inference of Gene Regulatory NetworksabstractModeling gene regulatory networks (GRNs) is essential for conceptualizing how genes are expressed and how they influence each other. Typically, a reverse engineering approach is employed; this strategy is effective in reproducing possible fitting models of GRNs. To use this strategy, however, two daunting tasks must be undertaken: one task is to optimize the accuracy of inferred network behaviors; and the other task is to designate valid biological topologies for target networks. Although existing studies have addressed these two tasks for years, few of the studies can satisfy both of the requirements simultaneously. To address these difficulties, we propose an integrative modeling framework that combines knowledge-based and data-driven input sources to construct biological topologies with their corresponding network behaviors. To validate the proposed approach, a real dataset collected from the cell cycle of the yeast S. cerevisiae is used. The results show that the proposed framework can successfully infer solutions that meet the requirements of both the network behaviors and biological structures. Therefore, the outcomes are exploitable for future in vivo experimental design. Yu-Ting Hsiao, Wei-Po Lee, Stefan Müller 0009, Christoph Flamm, Ivo L. Hofacker, Philipp Kügler |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2015 | Forna (force-directed RNA): Simple and effective online RNA secondary structure diagramsabstractMOTIVATION: The secondary structure of RNA is integral to the variety of functions it carries out in the cell and its depiction allows researchers to develop hypotheses about which nucleotides and base pairs are functionally relevant. Current approaches to visualizing secondary structure provide an adequate platform for the conversion of static text-based representations to 2D images, but are limited in their offer of interactivity as well as their ability to display larger structures, multiple structures and pseudoknotted structures. RESULTS: In this article, we present forna, a web-based tool for displaying RNA secondary structure which allows users to easily convert sequences and secondary structures to clean, concise and customizable visualizations. It supports, among other features, the simultaneous visualization of multiple structures, the display of pseudoknotted structures, the interactive editing of the displayed structures, and the automatic generation of secondary structure diagrams from PDB files. It requires no software installation apart from a modern web browser. AVAILABILITY AND IMPLEMENTATION: The web interface of forna is available at http://rna.tbi.univie.ac.at/forna while the source code is available on github at www.github.com/pkerpedjiev/forna. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peter Kerpedjiev, Stefan Hammer, Ivo L. Hofacker |
Bioinform. | 3 |
| 2015 | Model-Free RNA Sequence and Structure Alignment Informed by SHAPE Probing Reveals a Conserved Alternate Secondary Structure for 16S rRNAabstractDiscovery and characterization of functional RNA structures remains challenging due to deficiencies in de novo secondary structure modeling. Here we describe a dynamic programming approach for model-free sequence comparison that incorporates high-throughput chemical probing data. Based on SHAPE probing data alone, ribosomal RNAs (rRNAs) from three diverse organisms--the eubacteria E. coli and C. difficile and the archeon H. volcanii--could be aligned with accuracies comparable to alignments based on actual sequence identity. When both base sequence identity and chemical probing reactivities were considered together, accuracies improved further. Derived sequence alignments and chemical probing data from protein-free RNAs were then used as pseudo-free energy constraints to model consensus secondary structures for the 16S and 23S rRNAs. There are critical differences between these experimentally-informed models and currently accepted models, including in the functionally important neck and decoding regions of the 16S rRNA. We infer that the 16S rRNA has evolved to undergo large-scale changes in base pairing as part of ribosome function. As high-quality RNA probing data become widely available, structurally-informed sequence alignment will become broadly useful for de novo motif and function discovery. Christopher A. Lavender, Ronny Lorenz, Rita Tamayo, Ivo L. Hofacker, Kevin M. Weeks |
PLoS Comput. Biol. | 5 |
| 2015 | Product Grammars for Alignment and FoldingabstractWe develop a theory of algebraic operations over linear and context-free grammars that makes it possible to combine simple "atomic" grammars operating on single sequences into complex, multi-dimensional grammars. We demonstrate the utility of this framework by constructing the search spaces of complex alignment problems on multiple input sequences explicitly as algebraic expressions of very simple one-dimensional grammars. In particular, we provide a fully worked frameshift-aware, semiglobal DNA-protein alignment algorithm whose grammar is composed of products of small, atomic grammars. The compiler accompanying our theory makes it easy to experiment with the combination of multiple grammars and different operations. Composite grammars can be written out in L(A)T(E)X for documentation and as a guide to implementation of dynamic programming algorithms. An embedding in Haskell as a domain-specific language makes the theory directly accessible to writing and using grammar products without the detour of an external compiler. Software and supplemental files available here: http://www.bioinf. uni-leipzig.de/Software/gramprod/. Christian Höner zu Siederdissen, Ivo L. Hofacker, Peter F. Stadler |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | Basin Hopping Graph: a computational framework to characterize RNA folding landscapesabstractMOTIVATION: RNA folding is a complicated kinetic process. The minimum free energy structure provides only a static view of the most stable conformational state of the system. It is insufficient to give detailed insights into the dynamic behavior of RNAs. A sufficiently sophisticated analysis of the folding free energy landscape, however, can provide the relevant information. RESULTS: We introduce the Basin Hopping Graph (BHG) as a novel coarse-grained model of folding landscapes. Each vertex of the BHG is a local minimum, which represents the corresponding basin in the landscape. Its edges connect basins when the direct transitions between them are 'energetically favorable'. Edge weights endcode the corresponding saddle heights and thus measure the difficulties of these favorable transitions. BHGs can be approximated accurately and efficiently for RNA molecules well beyond the length range accessible to enumerative algorithms. AVAILABILITY AND IMPLEMENTATION: The algorithms described here are implemented in C++ as standalone programs. Its source code and supplemental material can be freely downloaded from http://www.tbi.univie.ac.at/bhg.html. Marcel Kucharík, Ivo L. Hofacker, Peter F. Stadler, Jing Qin 0006 |
Bioinform. | 2 |
| 2014 | Challenges in RNA virus bioinformaticsabstractMOTIVATION: Computer-assisted studies of structure, function and evolution of viruses remains a neglected area of research. The attention of bioinformaticians to this interesting and challenging field is far from commensurate with its medical and biotechnological importance. It is telling that out of >200 talks held at ISMB 2013, the largest international bioinformatics conference, only one presentation explicitly dealt with viruses. In contrast to many broad, established and well-organized bioinformatics communities (e.g. structural genomics, ontologies, next-generation sequencing, expression analysis), research groups focusing on viruses can probably be counted on the fingers of two hands. RESULTS: The purpose of this review is to increase awareness among bioinformatics researchers about the pressing needs and unsolved problems of computational virology. We focus primarily on RNA viruses that pose problems to many standard bioinformatics analyses owing to their compact genome organization, fast mutation rate and low evolutionary conservation. We provide an overview of tools and algorithms for handling viral sequencing data, detecting functionally important RNA structures, classifying viral proteins into families and investigating the origin and evolution of viruses. Manja Marz, Niko Beerenwinkel, Christian Drosten, Markus Fricke, Dmitrij Frishman, Ivo L. Hofacker, Dieter Hoffmann, Martin Middendorf, Thomas Rattei, Peter F. Stadler, Armin Töpfer |
Bioinform. | 6 |
| 2014 | TSSAR: TSS annotation regime for dRNA-seq dataabstractBACKGROUND: Differential RNA sequencing (dRNA-seq) is a high-throughput screening technique designed to examine the architecture of bacterial operons in general and the precise position of transcription start sites (TSS) in particular. Hitherto, dRNA-seq data were analyzed by visualizing the sequencing reads mapped to the reference genome and manually annotating reliable positions. This is very labor intensive and, due to the subjectivity, biased. RESULTS: Here, we present TSSAR, a tool for automated de novo TSS annotation from dRNA-seq data that respects the statistics of dRNA-seq libraries. TSSAR uses the premise that the number of sequencing reads starting at a certain genomic position within a transcriptional active region follows a Poisson distribution with a parameter that depends on the local strength of expression. The differences of two dRNA-seq library counts thus follow a Skellam distribution. This provides a statistical basis to identify significantly enriched primary transcripts.We assessed the performance by analyzing a publicly available dRNA-seq data set using TSSAR and two simple approaches that utilize user-defined score cutoffs. We evaluated the power of reproducing the manual TSS annotation. Furthermore, the same data set was used to reproduce 74 experimentally validated TSS in H. pylori from reliable techniques such as RACE or primer extension. Both analyses showed that TSSAR outperforms the static cutoff-dependent approaches. CONCLUSIONS: Having an automated and efficient tool for analyzing dRNA-seq data facilitates the use of the dRNA-seq technique and promotes its application to more sophisticated analysis. For instance, monitoring the plasticity and dynamics of the transcriptomal architecture triggered by different stimuli and growth conditions becomes possible.The main asset of a novel tool for dRNA-seq analysis that reaches out to a broad user community is usability. As such, we provide TSSAR both as intuitive RESTful Web service ( http://rna.tbi.univie.ac.at/TSSAR) together with a set of post-processing and analysis tools, as well as a stand-alone version for use in high-throughput dRNA-seq data analysis pipelines. Fabian Amman, Michael T. Wolfinger, Ronny Lorenz, Ivo L. Hofacker, Peter F. Stadler, Sven Findeiß |
BMC Bioinform. | 4 |
| 2013 | 2D Meets 4G: G-Quadruplexes in RNA Secondary Structure PredictionabstractG-quadruplexes are abundant locally stable structural elements in nucleic acids. The combinatorial theory of RNA structures and the dynamic programming algorithms for RNA secondary structure prediction are extended here to incorporate G-quadruplexes using a simple but plausible energy model. With preliminary energy parameters, we find that the overwhelming majority of putative quadruplex-forming sequences in the human genome are likely to fold into canonical secondary structures instead. Stable G-quadruplexes are strongly enriched, however, in the 5'UTR of protein coding mRNAs. Ronny Lorenz, Stephan H. Bernhart, Jing Qin 0006, Christian Höner zu Siederdissen, Andrea Tanzer, Fabian Amman, Ivo L. Hofacker, Peter F. Stadler |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2012 | Folding RNA/DNA hybrid duplexesabstractMOTIVATION: While there are numerous programs that can predict RNA or DNA secondary structures, a program that predicts RNA/DNA hetero-dimers is still missing. The lack of easy to use tools for predicting their structure may be in part responsible for the small number of reports of biologically relevant RNA/DNA hetero-dimers. RESULTS: We present here an extension to the widely used ViennaRNA Package (Lorenz et al., 2011) for the prediction of the structure of RNA/DNA hetero-dimers. AVAILABILITY: http://www.tbi.univie.ac.at/~ronny/RNA/vrna2.html CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronny Lorenz, Ivo L. Hofacker, Stephan H. Bernhart |
Bioinform. | 2 |
| 2011 | A folding algorithm for extended RNA secondary structuresabstractMOTIVATION: RNA secondary structure contains many non-canonical base pairs of different pair families. Successful prediction of these structural features leads to improved secondary structures with applications in tertiary structure prediction and simultaneous folding and alignment. RESULTS: We present a theoretical model capturing both RNA pair families and extended secondary structure motifs with shared nucleotides using 2-diagrams. We accompany this model with a number of programs for parameter optimization and structure prediction. AVAILABILITY: All sources (optimization routines, RNA folding, RNA evaluation, extended secondary structure visualization) are published under the GPLv3 and available at www.tbi.univie.ac.at/software/rnawolf/. Christian Höner zu Siederdissen, Stephan H. Bernhart, Peter F. Stadler, Ivo L. Hofacker |
Bioinform. | 4 |
| 2011 | Fast accessibility-based prediction of RNA-RNA interactionsabstractMOTIVATION: Currently, the best RNA-RNA interaction prediction tools are based on approaches that consider both the inter- and intramolecular interactions of hybridizing RNAs. While accurate, these methods are too slow and memory-hungry to be employed in genome-wide RNA target scans. Alternative methods neglecting intramolecular structures are fast enough for genome-wide applications, but are too inaccurate to be of much practical use. RESULTS: A new approach for RNA-RNA interaction was developed, with a prediction accuracy that is similar to that of algorithms that explicitly consider intramolecular structures, but running at least three orders of magnitude faster than RNAup. This is achieved by using a combination of precomputed accessibility profiles with an approximate energy model. This approach is implemented in the new version of RNAplex. The software also provides a variant using multiple sequences alignments as input, resulting in a further increase in specificity. AVAILABILITY: RNAplex is available at www.bioinf.uni-leipzig.de/Software/RNAplex. Hakim Tafer, Fabian Amman, Florian Eggenhofer, Peter F. Stadler, Ivo L. Hofacker |
Bioinform. | 5 |
| 2011 | From Structure Prediction to Genomic Screens for Novel Non-Coding RNAsabstractNon-coding RNAs (ncRNAs) are receiving more and more attention not only as an abundant class of genes, but also as regulatory structural elements (some located in mRNAs). A key feature of RNA function is its structure. Computational methods were developed early for folding and prediction of RNA structure with the aim of assisting in functional analysis. With the discovery of more and more ncRNAs, it has become clear that a large fraction of these are highly structured. Interestingly, a large part of the structure is comprised of regular Watson-Crick and GU wobble base pairs. This and the increased amount of available genomes have made it possible to employ structure-based methods for genomic screens. The field has moved from folding prediction of single sequences to computational screens for ncRNAs in genomic sequence using the RNA structure as the main characteristic feature. Whereas early methods focused on energy-directed folding of single sequences, comparative analysis based on structure preserving changes of base pairs has been efficient in improving accuracy, and today this constitutes a key component in genomic screens. Here, we cover the basic principles of RNA folding and touch upon some of the concepts in current methods that have been applied in genomic screens for de novo RNA structures in searches for novel ncRNA genes and regulatory RNA structure on mRNAs. We discuss the strengths and weaknesses of the different strategies and how they can complement each other. Jan Gorodkin, Ivo L. Hofacker |
PLoS Comput. Biol. | 2 |
| 2010 | Discriminatory power of RNA family modelsabstractMOTIVATION: RNA family models group nucleotide sequences that share a common biological function. These models can be used to find new sequences belonging to the same family. To succeed in this task, a model needs to exhibit high sensitivity as well as high specificity. As model construction is guided by a manual process, a number of problems can occur, such as the introduction of more than one model for the same family or poorly constructed models. We explore the Rfam database to discover such problems. RESULTS: Our main contribution is in the definition of the discriminatory power of RNA family models, together with a first algorithm for its computation. In addition, we present calculations across the whole Rfam database that show several families lacking high specificity when compared to other families. We give a list of these clusters of families and provide a tentative explanation. Our program can be used to: (i) make sure that new models are not equivalent to any model already present in the database; and (ii) new models are not simply submodels of existing families. AVAILABILITY: www.tbi.univie.ac.at/software/cmcompare/. The code is licensed under the GPLv3. Results for the whole Rfam database and supporting scripts are available together with the software. Christian Höner zu Siederdissen, Ivo L. Hofacker |
Bioinform. | 2 |
| 2010 | RNAsnoop: efficient target prediction for H/ACA snoRNAsabstractMOTIVATION: Small nucleolar RNAs are an abundant class of non-coding RNAs that guide chemical modifications of rRNAs, snRNAs and some mRNAs. In the case of many 'orphan' snoRNAs, the targeted nucleotides remain unknown, however. The box H/ACA subclass determines uridine residues that are to be converted into pseudouridines via specific complementary binding in a well-defined secondary structure configuration that is outside the scope of common RNA (co-)folding algorithms. RESULTS: RNAsnoop implements a dynamic programming algorithm that computes thermodynamically optimal H/ACA-RNA interactions in an efficient scanning variant. Complemented by an support vector machine (SVM)-based machine learning approach to distinguish true binding sites from spurious solutions and a system to evaluate comparative information, it presents an efficient and reliable tool for the prediction of H/ACA snoRNA target sites. We apply RNAsnoop to identify the snoRNAs that are responsible for several of the remaining 'orphan' pseudouridine modifications in human rRNAs, and we assign a target to one of the five orphan H/ACA snoRNAs in Drosophila. AVAILABILITY: The C source code of RNAsnoop is freely available at http://www.tbi.univie.ac.at/ -htafer/RNAsnoop Hakim Tafer, Stephanie Kehr, Jana Schor, Ivo L. Hofacker, Peter F. Stadler |
Bioinform. | 4 |
| 2010 | Hybridization thermodynamics of NimbleGen MicroarraysabstractBACKGROUND: While microarrays are the predominant method for gene expression profiling, probe signal variation is still an area of active research. Probe signal is sequence dependent and affected by probe-target binding strength and the competing formation of probe-probe dimers and secondary structures in probes and targets. RESULTS: We demonstrate the benefits of an improved model for microarray hybridization and assess the relative contributions of the probe-target binding strength and the different competing structures. Remarkably, specific and unspecific hybridization were apparently driven by different energetic contributions: For unspecific hybridization, the melting temperature Tm was the best predictor of signal variation. For specific hybridization, however, the effective interaction energy that fully considered competing structures was twice as powerful a predictor of probe signal variation. We show that this was largely due to the effects of secondary structures in the probe and target molecules. The predictive power of the strength of these intramolecular structures was already comparable to that of the melting temperature or the free energy of the probe-target duplex. CONCLUSIONS: This analysis illustrates the importance of considering both the effects of probe-target binding strength and the different competing structures. For specific hybridization, the secondary structures of probe and target molecules turn out to be at least as important as the probe-target binding strength for an understanding of the observed microarray signal intensities. Besides their relevance for the design of new arrays, our results demonstrate the value of improving thermodynamic models for the read-out and interpretation of microarray signals. Ulrike Mückstein, Germán G. Leparc, Alexandra Posekany, Ivo L. Hofacker, David P. Kreil |
BMC Bioinform. | 4 |
| 2008 | SnoReport: computational identification of snoRNAs with unknown targetsabstractUNLABELLED: Unlike tRNAs and microRNAs, both classes of snoRNAs, which direct two distinct types of chemical modifications of uracil residues, have proved to be surprisingly difficult to find in genomic sequences. Most computational approaches so far have explicitly used the fact that snoRNAs predominantly target ribosomal RNAs and spliceosomal RNAs. The target is specified by a short stretch of sequence complementarity between the snoRNA and its target. This sequence complementarity to known targets crucially contributes to sensitivity and specificity of snoRNA gene finding algorithms. The discovery of 'orphan' snoRNAs, which either have no known target, or which target ordinary protein-coding mRNAs, however, begs the question whether this class of 'housekeeping' non-coding RNAs is much more widespread and might have a diverse set of regulatory functions. In order to approach this question, we present here a combination of RNA secondary structure prediction and machine learning that is designed to recognize the two major classes of snoRNAs, box C/D and box H/ACA snoRNAs, among ncRNA candidate sequences. The snoReport approach deliberately avoids any usage of target information. We find that the combination of the conserved sequence boxes and secondary structure constraints as a pre-filter with SVM classifiers based on a small set of structural descriptors are sufficient for a reliable identification of snoRNAs. Tests of snoReport on data from several recent experimental surveys show that the approach is feasible; the application to a dataset from a large-scale comparative genomics survey for ncRNAs suggests that there are likely hundreds of previously undescribed 'orphan' snoRNAs still hidden in the human genome. AVAILABILITY: The snoReport software is implemented in ANSI C. The source code is available under the GNU Public License at http://www.bioinf.uni-leipzig.de/Software/snoReport. Jana Schor, Ivo L. Hofacker, Peter F. Stadler |
Bioinform. | 2 |
| 2008 | RNAplex: a fast tool for RNA-RNA interaction searchabstractMOTIVATION: Regulatory RNAs often unfold their action via RNA-RNA interaction. Transcriptional gene silencing by means of siRNAs and miRNA as well as snoRNA directed RNA editing rely on this mechanism. Additionally ncRNA regulation in bacteria is mainly based upon RNA duplex formation. Finding putative target sites for newly discovered ncRNAs is a lengthy task as tools for cofolding RNA molecules like RNAcofold and RNAup are too slow for genome-wide search. Tools like RNAhybrid that neglects intramolecular interactions have runtimes proportional to O(m x n), albeit with a large prefactor. Still in many cases the need for even faster methods exists. RESULTS: We present a new program, RNAplex, especially designed to quickly find possible hybridization sites for a query RNA in large RNA databases. RNAplex uses a slightly different energy model which reduces the computational time by a factor 10-27 compared to RNAhybrid. In addition a length penalty allows to focus the target search on short highly stable interactions. AVAILABILITY: RNAplex can be downloaded at http://www.tbi.univie.ac.at/~htafer/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hakim Tafer, Ivo L. Hofacker |
Bioinform. | 2 |
| 2008 | RNAalifold: improved consensus structure prediction for RNA alignmentsabstractBACKGROUND: The prediction of a consensus structure for a set of related RNAs is an important first step for subsequent analyses. RNAalifold, which computes the minimum energy structure that is simultaneously formed by a set of aligned sequences, is one of the oldest and most widely used tools for this task. In recent years, several alternative approaches have been advocated, pointing to several shortcomings of the original RNAalifold approach. RESULTS: We show that the accuracy of RNAalifold predictions can be improved substantially by introducing a different, more rational handling of alignment gaps, and by replacing the rather simplistic model of covariance scoring with more sophisticated RIBOSUM-like scoring matrices. These improvements are achieved without compromising the computational efficiency of the algorithm. We show here that the new version of RNAalifold not only outperforms the old one, but also several other tools recently developed, on different datasets. CONCLUSION: The new version of RNAalifold not only can replace the old one for almost any application but it is also competitive with other approaches including those based on SCFGs, maximum expected accuracy, or hierarchical nearest neighbor classifiers. Stephan H. Bernhart, Ivo L. Hofacker, Sebastian Will, Andreas R. Gruber, Peter F. Stadler |
BMC Bioinform. | 2 |
| 2008 | Strategies for measuring evolutionary conservation of RNA secondary structuresabstractBACKGROUND: Evolutionary conservation of RNA secondary structure is a typical feature of many functional non-coding RNAs. Since almost all of the available methods used for prediction and annotation of non-coding RNA genes rely on this evolutionary signature, accurate measures for structural conservation are essential. RESULTS: We systematically assessed the ability of various measures to detect conserved RNA structures in multiple sequence alignments. We tested three existing and eight novel strategies that are based on metrics of folding energies, metrics of single optimal structure predictions, and metrics of structure ensembles. We find that the folding energy based SCI score used in the RNAz program and a simple base-pair distance metric are by far the most accurate. The use of more complex metrics like for example tree editing does not improve performance. A variant of the SCI performed particularly well on highly conserved alignments and is thus a viable alternative when only little evolutionary information is available. Surprisingly, ensemble based methods that, in principle, could benefit from the additional information contained in sub-optimal structures, perform particularly poorly. As a general trend, we observed that methods that include a consensus structure prediction outperformed equivalent methods that only consider pairwise comparisons. CONCLUSION: Structural conservation can be measured accurately with relatively simple and intuitive metrics. They have the potential to form the basis of future RNA gene finders, that face new challenges like finding lineage specific structures or detecting mis-aligned sequences. Andreas R. Gruber, Stephan H. Bernhart, Ivo L. Hofacker, Stefan Washietl |
BMC Bioinform. | 3 |
| 2007 | Inferring Noncoding RNA Families and Classes by Means of Genome-Scale Structure-Based ClusteringabstractThe RFAM database defines families of ncRNAs by means of sequence similarities that are sufficient to establish homology. In some cases, such as microRNAs and box H/ACA snoRNAs, functional commonalities define classes of RNAs that are characterized by structural similarities, and typically consist of multiple RNA families. Recent advances in high-throughput transcriptomics and comparative genomics have produced very large sets of putative noncoding RNAs and regulatory RNA signals. For many of them, evidence for stabilizing selection acting on their secondary structures has been derived, and at least approximate models of their structures have been computed. The overwhelming majority of these hypothetical RNAs cannot be assigned to established families or classes. We present here a structure-based clustering approach that is capable of extracting putative RNA classes from genome-wide surveys for structured RNAs. The LocARNA (local alignment of RNA) tool implements a novel variant of the Sankoff algorithm that is sufficiently fast to deal with several thousand candidate sequences. The method is also robust against false positive predictions, i.e., a contamination of the input data with unstructured or nonconserved sequences. We have successfully tested the LocARNA-based clustering approach on the sequences of the RFAM-seed alignments. Furthermore, we have applied it to a previously published set of 3,332 predicted structured elements in the Ciona intestinalis genome (Missal K, Rose D, Stadler PF (2005) Noncoding RNAs in Ciona intestinalis. Bioinformatics 21 (Supplement 2): i77-i78). In addition to recovering, e.g., tRNAs as a structure-based class, the method identifies several RNA families, including microRNA and snoRNA candidates, and suggests several novel classes of ncRNAs for which to date no representative has been experimentally characterized. Sebastian Will, Kristin Reiche, Ivo L. Hofacker, Peter F. Stadler, Rolf Backofen |
PLoS Comput. Biol. | 3 |
| 2006 | Local RNA base pairing probabilities in large sequencesabstractSUMMARY: The genome-wide search for non-coding RNAs requires efficient methods to compute and compare local secondary structures. Since the exact boundaries of such putative transcripts are typically unknown, arbitrary sequence windows have to be used in practice. Here we present a method for robustly computing the probabilities of local base pairs from long RNA sequences independent of the exact positions of the sequence window. AVAILABILITY: The program RNAplfold is part of the Vienna RNA Package and can be downloaded from http://www.tbi.univie.ac.at/RNA/. Stephan H. Bernhart, Ivo L. Hofacker, Peter F. Stadler |
Bioinform. | 2 |
| 2006 | Memory efficient folding algorithms for circular RNA secondary structuresabstractBACKGROUND: A small class of RNA molecules, in particular the tiny genomes of viroids, are circular. Yet most structure prediction algorithms handle only linear RNAs. The most straightforward approach is to compute circular structures from 'internal' and 'external' substructures separated by a base pair. This is incompatible, however, with the memory-saving approach of the Vienna RNA Package which builds a linear RNA structure from shorter (internal) structures only. RESULT: Here we describe how circular secondary structures can be obtained without additional memory requirements as a kind of 'post-processing' of the linear structures. AVAILABILITY: The circular folding algorithm is implemented in the current version of the of RNAfold program of the Vienna RNA Package, which can be downloaded from http://www.tbi.univie.ac.at/RNA/ Ivo L. Hofacker, Peter F. Stadler |
Bioinform. | 1 |
| 2006 | Thermodynamics of RNA-RNA bindingabstractBACKGROUND: Reliable prediction of RNA-RNA binding energies is crucial, e.g. for the understanding on RNAi, microRNA-mRNA binding and antisense interactions. The thermodynamics of such RNA-RNA interactions can be understood as the sum of two energy contributions: (1) the energy necessary to 'open' the binding site and (2) the energy gained from hybridization. METHODS: We present an extension of the standard partition function approach to RNA secondary structures that computes the probabilities Pu[i, j] that a sequence interval [i, j] is unpaired. RESULTS: Comparison with experimental data shows that Pu[i, j] can be applied as a significant determinant of local target site accessibility for RNA interference (RNAi). Furthermore, these quantities can be used to rigorously determine binding free energies of short oligomers to large mRNA targets. The resource consumption is comparable with a single partition function computation for the large target molecule. We can show that RNAi efficiency correlates well with the binding energies of siRNAs to their respective mRNA target. AVAILABILITY: RNAup will be distributed as part of the Vienna RNA Package, www.tbi.univie.ac.at/~ivo/RNA/ Ulrike Mückstein, Hakim Tafer, Jörg Hackermüller, Stephan H. Bernhart, Peter F. Stadler, Ivo L. Hofacker |
Bioinform. | 6 |
| 2006 | Algebraic comparison of metabolic networks, phylogenetic inference, and metabolic innovationabstractBACKGROUND: Comparison of metabolic networks is typically performed based on the organisms' enzyme contents. This approach disregards functional replacements as well as orthologies that are misannotated. Direct comparison of the structure of metabolic networks can circumvent these problems. RESULTS: Metabolic networks are naturally represented as directed hypergraphs in such a way that metabolites are nodes and enzyme-catalyzed reactions form (hyper)edges. The familiar operations from set algebra (union, intersection, and difference) form a natural basis for both the pairwise comparison of networks and identification of distinct metabolic features of a set of algorithms. We report here on an implementation of this approach and its application to the procaryotes. CONCLUSION: We demonstrate that metabolic networks contain valuable phylogenetic information by comparing phylogenies obtained from network comparisons with 16S RNA phylogenies. The algebraic approach to metabolic networks is suitable to study metabolic innovations in two sets of organisms, free living microbes and Pyrococci, as well as obligate intracellular pathogens. Christian V. Forst, Christoph Flamm, Ivo L. Hofacker, Peter F. Stadler |
BMC Bioinform. | 3 |
| 2006 | Visualization of Barrier Tree SequencesabstractDynamical models that explain the formation of spatial structures of RNA molecules have reached a complexity that requires novel visualization methods that help to analyze the validity of these models. Here, we focus on the visualization of so-called folding landscapes of a growing RNA molecule. Folding landscapes describe the energy of a molecule as a function of its spatial configuration; thus they are huge and high dimensional. Their most salient features, however, are encapsulated by their so-called barrier tree that reflects the local minima and their connecting saddle points. For each length of the growing RNA chain there exists a folding landscape. We visualize the sequence of folding landscapes by an animation of the corresponding barrier trees. To generate the animation, we adapt the foresight layout with tolerance algorithm for general dynamic graph layout problems. Since it is very general, we give a detailed description of each phase: constructing a supergraph for the trees, layout of that supergraph using a modified DoT algorithm, and presentation techniques for the final animation. Christian Heine 0002, Gerik Scheuermann, Christoph Flamm, Ivo L. Hofacker, Peter F. Stadler |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2005 | Multiple sequence alignments of partially coding nucleic acid sequencesabstractBACKGROUND: High quality sequence alignments of RNA and DNA sequences are an important prerequisite for the comparative analysis of genomic sequence data. Nucleic acid sequences, however, exhibit a much larger sequence heterogeneity compared to their encoded protein sequences due to the redundancy of the genetic code. It is desirable, therefore, to make use of the amino acid sequence when aligning coding nucleic acid sequences. In many cases, however, only a part of the sequence of interest is translated. On the other hand, overlapping reading frames may encode multiple alternative proteins, possibly with intermittent non-coding parts. Examples are, in particular, RNA virus genomes. RESULTS: The standard scoring scheme for nucleic acid alignments can be extended to incorporate simultaneously information on translation products in one or more reading frames. Here we present a multiple alignment tool, codaln, that implements a combined nucleic acid plus amino acid scoring model for pairwise and progressive multiple alignments that allows arbitrary weighting for almost all scoring parameters. Resource requirements of codaln are comparable with those of standard tools such as ClustalW. CONCLUSION: We demonstrate the applicability of codaln to various biologically relevant types of sequences (bacteriophage Levivirus and Vertebrate Hox clusters) and show that the combination of nucleic acid and amino acid sequence information leads to improved alignments. These, in turn, increase the performance of analysis tools that depend strictly on good input alignments such as methods for detecting conserved RNA secondary structure elements. Roman R. Stocsits, Ivo L. Hofacker, Claudia Fried, Peter F. Stadler |
BMC Bioinform. | 2 |
| 2004 | Alignment of RNA base pairing probability matricesabstractMOTIVATION: Many classes of functional RNA molecules are characterized by highly conserved secondary structures but little detectable sequence similarity. Reliable multiple alignments can therefore be constructed only when the shared structural features are taken into account. Since multiple alignments are used as input for many subsequent methods of data analysis, structure-based alignments are an indispensable necessity in RNA bioinformatics. RESULTS: We present here a method to compute pairwise and progressive multiple alignments from the direct comparison of base pairing probability matrices. Instead of attempting to solve the folding and the alignment problem simultaneously as in the classical Sankoff's algorithm, we use McCaskill's approach to compute base pairing probability matrices which effectively incorporate the information on the energetics of each sequences. A novel, simplified variant of Sankoff's algorithms can then be employed to extract the maximum-weight common secondary structure and an associated alignment. AVAILABILITY: The programs pmcomp and pmmulti described in this contribution are implemented in Perl and can be downloaded together with the example datasets from http://www.tbi.univie.ac.at/RNA/PMcomp/. A web server is available at http://rna.tbi.univie.ac.at/cgi-bin/pmcgi.pl Ivo L. Hofacker, Stephan H. Bernhart, Peter F. Stadler |
Bioinform. | 1 |
| 2004 | Prediction of locally stable RNA secondary structures for genome-wide surveysabstractMOTIVATION: Recently novel classes of functional RNAs, most prominently the miRNAs have been discovered, strongly suggesting that further types of functional RNAs are still hidden in the recently completed genomic DNA sequences. Only few techniques are known, however, to survey genomes for such RNA genes. When sufficiently similar sequences are not available for comparative approaches the only known remedy is to search directly for structural features. RESULTS: We present here efficient algorithms for computing locally stable RNA structures at genome-wide scales. Both the minimum energy structure and the complete matrix of base pairing probabilities can be computed in theta(N x L2) time and theta(N + L2) memory in terms of the length N of the genome and the size L of the largest secondary structure motifs of interest. In practice, the 100 Mb of the complete genome of Caenorhabditis elegans can be folded within about half a day on a modern PC with a search depth of L = 100. This is sufficient example for a survey for miRNAs. AVAILABILITY: The software described in this contribution will be available for download at http://www.tbi.univie.ac.at/~ivo/RNA/ as part of the Vienna RNA Package. Ivo L. Hofacker, Barbara Priwitzer, Peter F. Stadler |
Bioinform. | 1 |
| 2004 | Conserved RNA secondary structures in viral genomes: a surveyabstractSUMMARY: The genomes of RNA viruses often carry conserved RNA structures that perform vital functions during the life cycle of the virus. Such structures can be detected using a combination of structure prediction and co-variation analysis. Here we present results from pilot studies on a variety of viral families performed during bioinformatics computer lab courses in past years. Ivo L. Hofacker, Peter F. Stadler, Roman R. Stocsits |
Bioinform. | 1 |
| 2004 | Prediction of Consensus RNA Secondary Structures Including PseudoknotsabstractMost functional RNA molecules have characteristic structures that are highly conserved in evolution. Many of them contain pseudoknots. Here, we present a method for computing the consensus structures including pseudoknots based on alignments of a few sequences. The algorithm combines thermodynamic and covariation information to assign scores to all possible base pairs, the base pairs are chosen with the help of the maximum weighted matching algorithm. We applied our algorithm to a number of different types of RNA known to contain pseudoknots. All pseudoknots were predicted correctly and more than 85 percent of the base pairs were identified. Christina Witwer, Ivo L. Hofacker, Peter F. Stadler |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 1998 | Combinatorics of RNA Secondary Structures
Ivo L. Hofacker, Peter Schuster 0002, Peter F. Stadler |
Discret. Appl. Math. | 1 |
| 1996 | Knowledge Discovery in RNA Sequence Families of HIV Using Scalable Computers
Ivo L. Hofacker, Martijn A. Huynen, Peter F. Stadler, Paul E. Stolorz |
KDD | 1 |