Mateo Gray

dblp:320/2489 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0001-7143-1367ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2026 PRISM: Partition-Function Decomposition into Structural Classes for Hierarchically Constrained RNA Pseudoknot Ensembles
abstract
While structure ensemble analysis became a valuable routinely applied tool for pseudoknot-free RNA, the extension to pseudoknots remains challenging due to the computational hardness of the general problem. The existing efficient algorithms for the computation of partition function with pseudoknots were still computationally expensive and were restricted to simple pseudoknots. This changed only with CParty, which computes pseudoknotted partition functions with the efficiency of pseudoknot-free folding. At its core, CParty follows the hierarchical folding hypothesis, such that ensemble structures can form pseudoknots only with a given input constraint structure. For an RNA sequence S and pseudoknot-free structure G, CParty limits the ensemble to "density-2" structures G∪ G' for a second, disjoint pseudoknot-free structure G'. We present PRISM that extends CParty from pure partition function calculation to full-fledged posterior probability analysis. By stochastic traceback through CParty’s dynamic programming matrices, it samples structures from the conditional Boltzmann ensemble. From estimated base pair probabilities, it generates ensemble representations, predicts centroid and maximum expected accuracy structures and calculates properties. In addition to position-specific summaries, PRISM maps sampled structures to RNA shapes, producing a posterior distribution over topological abstractions. This shape-level summary captures ensemble diversity even when a conserved pseudoknotted motif appears with shifted base-pair positions across samples. We validate PRISM in the pseudoknot-free limit, where it reproduces RNAFold quantities for minimum free energy, ensemble free energy, centroid expected distance, and maximum expected accuracy. We further show that stochastic traceback recovers Boltzmann structure probabilities and that sampling error decreases at the expected Monte Carlo rate while runtime grows linearly with the number of samples. Our case study demonstrate that RNA-shape summaries can reveal dominant pseudoknotted topologies that centroid decoding may miss. PRISM thus converts the CParty partition function into a practical framework for posterior decoding and topology-aware analysis of hierarchically constrained pseudoknotted RNA ensembles.
Mateo Gray, Sebastian Will, Hosna Jabbari
WABI1
2026 Spark: sparse hierarchical energy minimization for scalable prediction of RNA pseudoknots
abstract
MOTIVATION: The biological functions of RNAs are tightly connected to their specific RNA structures. As experimental techniques to determine high-accuracy structures are costly and time-consuming, computational prediction approaches became indispensable for biological RNA research; most notably, the prediction of minimum free energy secondary structures. Pseudoknots are prevalent, highly significant structural motifs, yet they are commonly ignored to achieve acceptable efficiency. Existing reliable pseudoknot prediction methods typically have prohibitive complexity. A route to fast scalable pseudoknot prediction was suggested with HFold following the hierarchical folding hypothesis. Recent successful sparsification of the CCJ pseudoknot prediction algorithm in Knotty promises a further boost by introducing this technique to hierarchical folding. RESULTS: We introduce Spark, a sparsified algorithm for predicting pseudoknotted RNA structures. Spark predicts exactly the same minimum-energy structures as its predecessor HFold in the accurate HotKnots 2.0 energy model for pseudoknots. While sparsification maintains exact energy minimization and theoretical complexity, it strongly improves the time and space consumption over HFold. We benchmarked the performance of Spark against HFold and, as a pseudoknot-free baseline, RNAfold. Compared with HFold, Spark substantially reduces both run time and memory usage, while achieving run times close to RNAfold. Across all tested sequence lengths, Spark used the least memory and consistently ran faster than HFold. CONCLUSION: Combining sparsification and hierarchical folding in Spark results in an remarkably fast and memory-efficient tool for the accurate prediction of pseudoknotted RNA structures. Consequently, Spark practically enables pseudoknot prediction in large scale and even for very long RNA sequences. AVAILABILITY: Spark software is available on Github (https://github.com/TheCOBRALab/Spark), with a permanent archive of the software and results deposited on Zenodo (https://doi.org/10.5281/zenodo.19073315).
Mateo Gray, Sebastian Will, Hosna Jabbari
Bioinform.1
2025 Spark: Sparsified Hierarchical Energy Minimization of RNA Pseudoknots
Mateo Gray, Sebastian Will, Hosna Jabbari
WABI1
2025 CParty: hierarchically constrained partition function of RNA pseudoknots
abstract
MOTIVATION: Biologically relevant RNA secondary structures are routinely predicted by efficient dynamic programming algorithms that minimize their free energy. Starting from such algorithms, one can devise partition function algorithms, which enable stochastic perspectives on RNA structure ensembles. As the most prominent example, McCaskill's partition function algorithm is derived from pseudoknot-free energy minimization. While this algorithm became hugely successful for the analysis of pseudoknot-free RNA structure ensembles, as of yet there exists only one pseudoknotted partition function implementation, which covers only simple pseudoknots and comes with a borderline-prohibitive complexity of O(n5) in the RNA length n. RESULTS: Here, we develop a partition function algorithm corresponding to the hierarchical pseudoknot prediction of HFold, which performs exact optimization in a realistic pseudoknot energy model. In consequence, our algorithm CParty carries over HFold's advantages over classical pseudoknot prediction in characterizing the Boltzmann ensemble at equilibrium. Given an RNA sequence S and a pseudoknot-free structure G, CParty computes the partition function over all possibly pseudoknotted density-2 structures G∪G' of S that extend the fixed G by a disjoint pseudoknot-free structure G'. Thus, CParty follows the common hypothesis of hierarchical pseudoknot formation, where pseudoknots form as tertiary contacts only after a first pseudoknot-free "core" G and we call the computed partition function hierarchically constrained (by G). Like HFold, the dynamic programming algorithm CParty is very efficient, achieving the low complexity of the pseudoknot-free algorithm, i.e. cubic time and quadratic space. Finally, by computing pseudoknotted ensemble energies, we unveil kinetics features of a therapeutic target in SARS-CoV-2. AVAILABILITY AND IMPLEMENTATION: CParty is available at https://github.com/HosnaJabbari/CParty.
Mateo Gray, Luke Trinity, Ulrike Stege, Yann Ponty, Sebastian Will, Hosna Jabbari
Bioinform.1
2023 SparseRNAFolD: Sparse RNA Pseudoknot-Free Folding Including Dangles
abstract
Motivation. Computational RNA secondary structure prediction by free energy minimization is indispensable for analyzing structural RNAs and their interactions. These methods find the structure with the minimum free energy (MFE) among exponentially many possible structures and have a restrictive time and space complexity (O(n³) time and O(n²) space for pseudoknot-free structures) for longer RNA sequences. Furthermore, accurate free energy calculations, including dangles contributions can be difficult and costly to implement, particularly when optimizing for time and space requirements. Results. Here we introduce a fast and efficient sparsified MFE pseudoknot-free structure prediction algorithm, SparseRNAFolD, that utilizes an accurate energy model that accounts for dangle contributions. While the sparsification technique was previously employed to improve the time and space complexity of a pseudoknot-free structure prediction method with a realistic energy model, SparseMFEFold, it was not extended to include dangle contributions due to the complexity of computation. This may come at the cost of prediction accuracy. In this work, we compare three different sparsified implementations for dangles contributions and provide pros and cons of each method. As well, we compare our algorithm to LinearFold, a linear time and space algorithm, where we find that in practice, SparseRNAFolD has lower memory consumption across all lengths of sequence and a faster time for lengths up to 1000 bases. Conclusion. Our SparseRNAFolD algorithm is an MFE-based algorithm that guarantees optimality of result and employs the most general energy model, including dangle contributions. We provide a basis for applying dangles to sparsified recursion in a pseudoknot-free model that has the ability to be extended to pseudoknots.
Mateo Gray, Sebastian Will, Hosna Jabbari
WABI1
2022 KnotAli: informed energy minimization through the use of evolutionary information
abstract
BACKGROUND: Improving the prediction of structures, especially those containing pseudoknots (structures with crossing base pairs) is an ongoing challenge. Homology-based methods utilize structural similarities within a family to predict the structure. However, their prediction is limited to the consensus structure, and by the quality of the alignment. Minimum free energy (MFE) based methods, on the other hand, do not rely on familial information and can predict structures of novel RNA molecules. Their prediction normally suffers from inaccuracies due to their underlying energy parameters. RESULTS: We present a new method for prediction of RNA pseudoknotted secondary structures that combines the strengths of MFE prediction and alignment-based methods. KnotAli takes a multiple RNA sequence alignment as input and uses covariation and thermodynamic energy minimization to predict possibly pseudoknotted secondary structures for each individual sequence in the alignment. We compared KnotAli's performance to that of three other alignment-based programs, two that can handle pseudoknotted structures and one control, on a large data set of 3034 RNA sequences with varying lengths and levels of sequence conservation from 10 families with pseudoknotted and pseudoknot-free reference structures. We produced sequence alignments for each family using two well-known sequence aligners (MUSCLE and MAFFT). CONCLUSIONS: We found KnotAli's performance to be superior in 6 of the 10 families for MUSCLE and 7 of the 10 for MAFFT. While both KnotAli and Cacofold use background noise correction strategies, we found KnotAli's predictions to be less dependent on the alignment quality. KnotAli can be found online at the Zenodo image: https://doi.org/10.5281/zenodo.5794719.
Mateo Gray, Sean Chester, Hosna Jabbari
BMC Bioinform.1