Tomasz Zok

dblp:12/10953 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-4103-9238ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2025 RNAtive to recognize native-like structure in a set of RNA 3D models
abstract
MOTIVATION: Most widely used methods for evaluating RNA 3D structure models require experimental reference structures, which restricts their use for novel RNAs. They also often overlook recurrent structural features shared across multiple predictions of the same sequence. Although consensus approaches have proven effective in RNA sequence analysis and evolutionary studies, no existing tool applies these principles to evaluate ensembles of 3D models. This gap hampers the identification of native-like folds in computational predictions, particularly as AI-driven methods become increasingly prevalent. RESULTS: This paper presents RNAtive, the first computational tool to apply consensus-derived secondary structures for reference-free evaluation of RNA 3D models. RNAtive aggregates recurrent base-pairing and stacking interactions across ensembles of predicted 3D structures to construct a consensus secondary structure. It introduces a novel conditionally weighted consensus mode that treats interaction networks as fuzzy sets and uniquely allows integration of user-defined 2D structural constraints, enabling evaluation guided by experimental data. Input RNA models are ranked using two adapted binary-classification-based scores. Benchmarking against CASP15 competition data shows that models consistent with the consensus exhibit native-like structural features. The RNAtive web server offers an intuitive platform for comparing and prioritizing RNA 3D predictions, providing a scalable solution to address the variability inherent in deep learning and fragment-assembly methods. By bridging consensus principles with 3D structural analysis, RNAtive advances the exploration of RNA conformational landscapes and has potential applications in fields like therapeutic RNA design. AVAILABILITY: RNAtive is a freely accessible web server with a modern, user-friendly interface, available for scientific, educational, and commercial use at https://rnative.cs.put.poznan.pl/.
Jan Pielesiak, Maciej Antczak, Marta Szachniuk, Tomasz Zok
Bioinform.4
2024 Knotted artifacts in predicted 3D RNA structures
abstract
Unlike proteins, RNAs deposited in the Protein Data Bank do not contain topological knots. Recently, admittedly, the first trefoil knot and some lasso-type conformations have been found in experimental RNA structures, but these are still exceptional cases. Meanwhile, algorithms predicting 3D RNA models have happened to form knotted structures not so rarely. Interestingly, machine learning-based predictors seem to be more prone to generate knotted RNA folds than traditional methods. A similar situation is observed for the entanglements of structural elements. In this paper, we analyze all models submitted to the CASP15 competition in the 3D RNA structure prediction category. We show what types of topological knots and structure element entanglements appear in the submitted models and highlight what methods are behind the generation of such conformations. We also study the structural aspect of susceptibility to entanglement. We suggest that predictors take care of an evaluation of RNA models to avoid publishing structures with artifacts, such as unusual entanglements, that result from hallucinations of predictive algorithms.
Bartosz Ambrozy Gren, Maciej Antczak, Tomasz Zok, Joanna I. Sulkowska, Marta Szachniuk
PLoS Comput. Biol.3
2024 RNAtango: Analysing and comparing RNA 3D structures via torsional angles
abstract
RNA molecules, essential for viruses and living organisms, derive their pivotal functions from intricate 3D structures. To understand these structures, one can analyze torsion and pseudo-torsion angles, which describe rotations around bonds, whether real or virtual, thus capturing the RNA conformational flexibility. Such an analysis has been made possible by RNAtango, a web server introduced in this paper, that provides a trigonometric perspective on RNA 3D structures, giving insights into the variability of examined models and their alignment with reference targets. RNAtango offers comprehensive tools for calculating torsion and pseudo-torsion angles, generating angle statistics, comparing RNA structures based on backbone torsions, and assessing local and global structural similarities using trigonometric functions and angle measures. The system operates in three scenarios: single model analysis, model-versus-target comparison, and model-versus-model comparison, with results output in text and graphical formats. Compatible with all modern web browsers, RNAtango is accessible freely along with the source code. It supports researchers in accurately assessing structural similarities, which contributes to the precision and efficiency of RNA modeling.
Marta Mackowiak, Bartosz Adamczyk, Marta Szachniuk, Tomasz Zok
PLoS Comput. Biol.4
2022 Characterizing domain-specific open educational resources by linking ISCB Communities of Special Interest to Wikipedia
abstract
MOTIVATION: Wikipedia is one of the most important channels for the public communication of science and is frequently accessed as an educational resource in computational biology. Joint efforts between the International Society for Computational Biology (ISCB) and the Computational Biology taskforce of WikiProject Molecular Biology (a group of expert Wikipedia editors) have considerably improved computational biology representation on Wikipedia in recent years. However, there is still an urgent need for further improvement in quality, especially when compared to related scientific fields such as genetics and medicine. Facilitating involvement of members from ISCB Communities of Special Interest (COSIs) would improve a vital open education resource in computational biology, additionally allowing COSIs to provide a quality educational resource highly specific to their subfield. RESULTS: We generate a list of around 1500 English Wikipedia articles relating to computational biology and describe the development of a binary COSI-Article matrix, linking COSIs to relevant articles and thereby defining domain-specific open educational resources. Our analysis of the COSI-Article matrix data provides a quantitative assessment of computational biology representation on Wikipedia against other fields and at a COSI-specific level. Furthermore, we conducted similarity analysis and subsequent clustering of COSI-Article data to provide insight into potential relationships between COSIs. Finally, based on our analysis, we suggest courses of action to improve the quality of computational biology representation on Wikipedia.
Alastair M. Kilpatrick, Farzana Rahman, Audra Anjum, Sayane Shome, K. M. Salim Andalib, Shrabonti Banik, Sanjana F. Chowdhury, Peter Coombe, Yesid Cuesta Astroz, J. Maxwell Douglas, Pradeep Eranti, Aleyna D. Kiran, Sachendra Kumar, Hyeri Lim, Valentina Lorenzi, Tiago Lubiana, Sakib Mahmud, Rafael Puche, Agnieszka Rybarczyk, Syed Muktadir Al Sium, David Twesigomwe, Tomasz Zok, Christine A. Orengo, Iddo Friedberg, Janet Kelso, Lonnie R. Welch
Bioinform.22
2022 RNAloops: a database of RNA multiloops
abstract
MOTIVATION: Knowledge of the 3D structure of RNA supports discovering its functions and is crucial for designing drugs and modern therapeutic solutions. Thus, much attention is devoted to experimental determination and computational prediction targeting the global fold of RNA and its local substructures. The latter include multi-branched loops-functionally significant elements that highly affect the spatial shape of the entire molecule. Unfortunately, their computational modeling constitutes a weak point of structural bioinformatics. A remedy for this is in collecting these motifs and analyzing their features. RESULTS: RNAloops is a self-updating database that stores multi-branched loops identified in the PDB-deposited RNA structures. A description of each loop includes angular data-planar and Euler angles computed between pairs of adjacent helices to allow studying their mutual arrangement in space. The system enables search and analysis of multiloops, presents their structure details numerically and visually, and computes data statistics. AVAILABILITY AND IMPLEMENTATION: RNAloops is freely accessible at https://rnaloops.cs.put.poznan.pl. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jakub Wiedemann, Jacek Kaczor, Maciej Milostan, Tomasz Zok, Jacek Blazewicz, Marta Szachniuk, Maciej Antczak
Bioinform.4
2022 DrawTetrado to create layer diagrams of G4 structures
abstract
MOTIVATION: Quadruplexes are specific 3D structures found in nucleic acids. Due to the exceptional properties of these motifs, their exploration with the general-purpose bioinformatics methods can be problematic or insufficient. The same applies to visualizing their structure. A hand-drawn layer diagram is the most common way to represent the quadruplex anatomy. No molecular visualization software generates such a structural model based on atomic coordinates. RESULTS: DrawTetrado is an open-source Python program for automated visualization targeting the structures of quadruplexes and G4-helices. It generates static layer diagrams that represent structural data in a pseudo-3D perspective. The possibility to set color schemes, nucleotide labels, inter-element distances or angle of view allows for easy customization of the output drawing. AVAILABILITY AND IMPLEMENTATION: The program is available under the MIT license at https://github.com/RNApolis/drawtetrado.
Michal Zurkowski, Tomasz Zok, Marta Szachniuk
Bioinform.2
2021 BioCommons: a robust java library for RNA structural bioinformatics
abstract
MOTIVATION: Biomolecular structures come in multiple representations and diverse data formats. Their incompatibility with the requirements of data analysis programs significantly hinders the analytics and the creation of new structure-oriented bioinformatic tools. Therefore, the need for robust libraries of data processing functions is still growing. RESULTS: BioCommons is an open-source, Java library for structural bioinformatics. It contains many functions working with the 2D and 3D structures of biomolecules, with a particular emphasis on RNA. AVAILABILITY AND IMPLEMENTATION: The library is available in Maven Central Repository and its source code is hosted on GitHub: https://github.com/tzok/BioCommons. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tomasz Zok
Bioinform.1
2020 Topology-based classification of tetrads and quadruplex structures
abstract
MOTIVATION: Quadruplexes attract the attention of researchers from many fields of bio-science. Due to a specific structure, these tertiary motifs are involved in various biological processes. They are also promising therapeutic targets in many strategies of drug development, including anticancer and neurological disease treatment. The uniqueness and diversity of their forms cause that quadruplexes show great potential in novel biological applications. The existing approaches for quadruplex analysis are based on sequence or 3D structure features and address canonical motifs only. RESULTS: In our study, we analyzed tetrads and quadruplexes contained in nucleic acid molecules deposited in Protein Data Bank. Focusing on their secondary structure topology, we adjusted its graphical diagram and proposed new dot-bracket and arc representations. We defined the novel classification of these motifs. It can handle both canonical and non-canonical cases. Based on this new taxonomy, we implemented a method that automatically recognizes the types of tetrads and quadruplexes occurring as unimolecular structures. Finally, we conducted a statistical analysis of these motifs found in experimentally determined nucleic acid structures in relation to the new classification. AVAILABILITY AND IMPLEMENTATION: https://github.com/tzok/eltetrado/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mariusz Popenda, Joanna Miskiewicz, Joanna Sarzynska, Tomasz Zok, Marta Szachniuk
Bioinform.4
2020 ElTetrado: a tool for identification and classification of tetrads and quadruplexes
abstract
BACKGROUND: Quadruplexes are specific structure motifs occurring, e.g., in telomeres and transcriptional regulatory regions. Recent discoveries confirmed their importance in biomedicine and led to an intensified examination of their properties. So far, the study of these motifs has focused mainly on the sequence and the tertiary structure, and concerned canonical structures only. Whereas, more and more non-canonical quadruplex motifs are being discovered. RESULTS: Here, we present ElTetrado, a software that identifies quadruplexes (composed of guanine- and other nucleobase-containing tetrads) in nucleic acid structures and classifies them according to the recently introduced ONZ taxonomy. The categorization is based on the secondary structure topology of quadruplexes and their component tetrads. It supports the analysis of canonical and non-canonical motifs. Besides the class recognition, ElTetrado prepares a dot-bracket and graphical representations of the secondary structure, which reflect the specificity of the quadruplex's structure topology. It is implemented as a freely available, standalone application, available at https://github.com/tzok/eltetrado. CONCLUSIONS: The proposed software tool allows to identify and classify tetrads and quadruplexes based on the topology of their secondary structures. It complements existing approaches focusing on the sequence and 3D structure.
Tomasz Zok, Mariusz Popenda, Marta Szachniuk
BMC Bioinform.1
2019 RNAvista: a webserver to assess RNA secondary structures with non-canonical base pairs
abstract
Motivation: In the study of 3D RNA structure, information about non-canonical interactions between nucleobases is increasingly important. Specialized databases support investigation of this issue based on experimental data, and several programs can annotate non-canonical base pairs in the RNA 3D structure. However, predicting the extended RNA secondary structure which describes both canonical and non-canonical interactions remains difficult. Results: Here, we present RNAvista that allows predicting an extended RNA secondary structure from sequence or from the list enumerating canonical base pairs only. RNAvista is implemented as a publicly available webserver with user-friendly interface. It runs on all major web browsers. Availability and implementation: http://rnavista.cs.put.poznan.pl.
Maciej Antczak, Marcin Zablocki, Tomasz Zok, Agnieszka Rybarczyk, Jacek Blazewicz, Marta Szachniuk
Bioinform.3
2018 New algorithms to represent complex pseudoknotted RNA structures in dot-bracket notation
abstract
Motivation: Understanding the formation, architecture and roles of pseudoknots in RNA structures are one of the most difficult challenges in RNA computational biology and structural bioinformatics. Methods predicting pseudoknots typically perform this with poor accuracy, often despite experimental data incorporation. Existing bioinformatic approaches differ in terms of pseudoknots' recognition and revealing their nature. A few ways of pseudoknot classification exist, most common ones refer to a genus or order. Following the latter one, we propose new algorithms that identify pseudoknots in RNA structure provided in BPSEQ format, determine their order and encode in dot-bracket-letter notation. The proposed encoding aims to illustrate the hierarchy of RNA folding. Results: New algorithms are based on dynamic programming and hybrid (combining exhaustive search and random walk) approaches. They evolved from elementary algorithm implemented within the workflow of RNA FRABASE 1.0, our database of RNA structure fragments. They use different scoring functions to rank dissimilar dot-bracket representations of RNA structure. Computational experiments show an advantage of new methods over the others, especially for large RNA structures. Availability and implementation: Presented algorithms have been implemented as new functionality of RNApdbee webserver and are ready to use at http://rnapdbee.cs.put.poznan.pl. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Maciej Antczak, Mariusz Popenda, Tomasz Zok, Michal Zurkowski, Ryszard W. Adamiak, Marta Szachniuk
Bioinform.3
2018 RNAfitme: a webserver for modeling nucleobase and nucleoside residue conformation in fixed-backbone RNA structures
abstract
BACKGROUND: Computational RNA 3D structure prediction and modeling are rising as complementary approaches to high-resolution experimental techniques for structure determination. They often apply to substitute or complement them. Recently, researchers' interests have directed towards in silico methods to fit, remodel and refine RNA tertiary structure models. Their power lies in a problem-specific exploration of RNA conformational space and efficient optimization procedures. The aim is to improve the accuracy of models obtained either computationally or experimentally. RESULTS: Here, we present RNAfitme, a versatile webserver tool for remodeling of nucleobase- and nucleoside residue conformations in the fixed-backbone RNA 3D structures. Our approach makes use of dedicated libraries that define RNA conformational space. They have been built upon torsional angle characteristics of PDB-deposited RNA structures. RNAfitme can be applied to reconstruct full-atom model of RNA from its backbone; remodel user-selected nucleobase/nucleoside residues in a given RNA structure; predict RNA 3D structure based on the sequence and the template of a homologous molecule of the same size; refine RNA 3D model by reducing steric clashes indicated during structure quality assessment. RNAfitme is a publicly available tool with an intuitive interface. It is freely accessible at http://rnafitme.cs.put.poznan.pl/ CONCLUSIONS: RNAfitme has been applied in various RNA 3D remodeling scenarios for several types of input data. Computational experiments proved its efficiency, accuracy, and usefulness in the processing of RNAs of any size. Fidelity of RNAfitme predictions has been thoroughly tested for RNA 3D structures determined experimentally and modeled in silico.
Maciej Antczak, Tomasz Zok, Maciej Osowiecki, Mariusz Popenda, Ryszard W. Adamiak, Marta Szachniuk
BMC Bioinform.2
2018 INDIGO-DataCloud: a Platform to Facilitate Seamless Access to E-Infrastructures
abstract
This paper describes the achievements of the H2020 project INDIGO-DataCloud. The project has provided e-infrastructures with tools, applications and cloud framework enhancements to manage the demanding requirements of scientific communities, either locally or through enhanced interfaces. The middleware developed allows to federate hybrid resources, to easily write, port and run scientific applications to the cloud. In particular, we have extended existing PaaS (Platform as a Service) solutions, allowing public and private e-infrastructures, including those provided by EGI, EUDAT, and Helix Nebula, to integrate their existing services and make them available through AAI services compliant with GEANT interfederation policies, thus guaranteeing transparency and trust in the provisioning of such services. Our middleware facilitates the execution of applications using containers on Cloud and Grid based infrastructures, as well as on HPC clusters. Our developments are freely downloadable as open source components, and are already being integrated into many scientific applications.
Davide Salomoni, Isabel Campos Plasencia, Luciano Gaido, Jesús E. Marco de Lucas, P. Solagna, Jorge Gomes 0001, Ludek Matyska, P. Fuhrman, Marcus Hardt, Giacinto Donvito, Lukasz Dutka, Marcin Plóciennik, Roberto Barbera, Ignacio Blanquer, Andrea Ceccanti, Eva Cetinic, Mário David, Doina Cristina Duma, Álvaro López García, Germán Moltó, Pablo Orviz Fernández, Zdenek Sustr, Matthew Viljoen, Fernando Aguilar, Marica Antonacci, Lucio Angelo Antonelli, Stefano Bagnasco, A. Bonving, Riccardo Bruno, Alessandro Costa, Davor Davidovic, Benjamin Ertl, Marco Fargetta, Sandro Fiore, S. Gallozzi, Z. Kurkcuoglu, Lara Lloret Iglesias, J. Martins, Alessandra Nuzzo, Paola Nassisi, Cosimo Palazzo, João Murta Pina, Eva Sciacca, Daniele Spiga, Marco Antonio Tangaro, Michal Urbaniak, Sara Vallero, Bas Wegh, Valentina Zaccolo, Federico Zambelli, Tomasz Zok
J. Grid Comput.53
2017 LCS-TA to identify similar fragments in RNA 3D structures
abstract
BACKGROUND: In modern structural bioinformatics, comparison of molecular structures aimed to identify and assess similarities and differences between them is one of the most commonly performed procedures. It gives the basis for evaluation of in silico predicted models. It constitutes the preliminary step in searching for structural motifs. In particular, it supports tracing the molecular evolution. Faced with an ever-increasing amount of available structural data, researchers need a range of methods enabling comparative analysis of the structures from either global or local perspective. RESULTS: Herein, we present a new, superposition-independent method which processes pairs of RNA 3D structures to identify their local similarities. The similarity is considered in the context of structure bending and bonds' rotation which are described by torsion angles. In the analyzed RNA structures, the method finds the longest continuous segments that show similar torsion within a user-defined threshold. The length of the segment is provided as local similarity measure. The method has been implemented as LCS-TA algorithm (Longest Continuous Segments in Torsion Angle space) and is incorporated into our MCQ4Structures application, freely available for download from http://www.cs.put.poznan.pl/tzok/mcq/ . CONCLUSIONS: The presented approach ties torsion-angle-based method of structure analysis with the idea of local similarity identification by handling continuous 3D structure segments. The first method, implemented in MCQ4Structures, has been successfully utilized in RNA-Puzzles initiative. The second one, originally applied in Euclidean space, is a component of LGA (Local-Global Alignment) algorithm commonly used in assessing protein models submitted to CASP. This unique combination of concepts implemented in LCS-TA provides a new perspective on structure quality assessment in local and quantitative aspect. A series of computational experiments show the first results of applying our method to comparison of RNA 3D models. LCS-TA can be used for identifying strengths and weaknesses in the prediction of RNA tertiary structures.
Jakub Wiedemann, Tomasz Zok, Maciej Milostan, Marta Szachniuk
BMC Bioinform.2
2016 Distributed and cloud-based multi-model analytics experiments on large volumes of climate change data in the earth system grid federation eco-system
abstract
A case study on climate models intercomparison data analysis addressing several classes of multi-model experiments is being implemented in the context of the EU H2020 INDIGO-DataCloud project. Such experiments require the availability of large amount of data (multi-terabyte order) related to the output of several climate models simulations as well as the exploitation of scientific data management tools for large-scale data analytics. More specifically, the paper discusses in detail a use case on precipitation trend analysis in terms of requirements, architectural design solution, and infrastructural implementation. The experiment has been tested and validated on CMIP5 datasets, in the context of a large scale distributed testbed across EU and US involving three ESGF sites (LLNL, ORNL, and CMCC) and one central orchestrator site (PSNC).
Sandro Fiore, Marcin Plóciennik, Charles M. Doutriaux, Cosimo Palazzo, Jason Boutte, Tomasz Zok, Donatello Elia, Michal Owsiak, Alessandro D'Anca, Z. Shaheen, Riccardo Bruno, Marco Fargetta, Miguel Caballer, Germán Moltó, Ignacio Blanquer, Roberto Barbera, Mário David, Giacinto Donvito, Dean N. Williams, Valentine Anantharaj, Davide Salomoni, Giovanni Aloisio
IEEE BigData6
2015 New in silico approach to assessing RNA secondary structures with non-canonical base pairs
abstract
BACKGROUND: The function of RNA is strongly dependent on its structure, so an appropriate recognition of this structure, on every level of organization, is of great importance. One particular concern is the assessment of base-base interactions, described as the secondary structure, the knowledge of which greatly facilitates an interpretation of RNA function and allows for structure analysis on the tertiary level. The RNA secondary structure can be predicted from a sequence using in silico methods often adjusted with experimental data, or assessed from 3D structure atom coordinates. Computational approaches typically consider only canonical, Watson-Crick and wobble base pairs. Handling of non-canonical interactions, important for a full description of RNA structure, is still very difficult. RESULTS: We introduce our novel approach to assessing an extended RNA secondary structure, which characterizes both canonical and non-canonical base pairs, along with their type classification. It is based on predicting the RNA 3D structure from a user-provided sequence or a secondary structure that only describes canonical base pairs, and then deriving the extended secondary structure from atom coordinates. In our example implementation, this was achieved by integrating the functionality of two fully automated, high fidelity methods in a computational pipeline: RNAComposer for the 3D RNA structure prediction and RNApdbee for base-pair annotation. CONCLUSIONS: The presented methodology ties together existing applications for RNA 3D structure prediction and base-pair annotation. The example performance, applying RNAComposer and RNApdbee, reveals better accuracy in non-canonical base pair assessment than the compared methods that directly predict RNA secondary structure.
Agnieszka Rybarczyk, Natalia Szostak, Maciej Antczak, Tomasz Zok, Mariusz Popenda, Ryszard W. Adamiak, Jacek Blazewicz, Marta Szachniuk
BMC Bioinform.4
2013 Approaches to Distributed Execution of Scientific Workflows in Kepler
abstract
The Kepler scientific workflow system enables creation, execution and sharing of workflows across a broad range of scientific and engineering disciplines while also facilitating remote and distributed execution of workflows. In this paper, we present
Marcin Plóciennik, Tomasz Zok, Ilkay Altintas, Jianwu Wang 0001, Daniel Crawl, David Abramson 0001, Frederic Imbeaux, Bernard Guillerminet, Marcos López-Caniego, Isabel Campos Plasencia, Wojciech Pych, Pawel Ciecielag, Bartek Palak, Michal Owsiak, Yann Frauel
Fundam. Informaticae2