David Hoksza

dblp:97/515 · DBLP profile ↗
← Back
48ranked-venue papers
14as first author
16since 2021 · last 2025
0000-0003-4679-0557ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 41 · 14 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Beyond Exact Matches: Near-Hit Scoring for Protein Binding Site Prediction with Protein Language Models
abstract
Deep understanding of protein behavior and its role in particular biological processes is a task crucial for many real-world applications, such as bioengineering or drug design. Recent advances in protein informatics, particularly Protein Language Models (PLMs), have transformed the field by capturing rich evolutionary, physicochemical, and structural features into high-dimensional embeddings. In this work, we investigate the application of PLMs to residue-level ligand-binding site prediction. Specifically, we show that standard evaluation metrics, which apply a hard spatial cutoff, systematically penalize biologically plausible “near-miss” predictions inherent to PLM embeddings. To address this, we introduce near-hit scoring, a distance-weighted evaluation metric that softens the boundary between correct and incorrect predictions by accounting for sequential proximity to annotated sites. Applying this framework to state-of-the-art PLM-based predictors on the LIGYSIS benchmark demonstrates that near-hit scoring provides a fairer and more informative assessment of binding site prediction performance.
Yana Podlesna, Vít Skrhák, David Hoksza
BIBM3
2025 Hybrid protein-ligand binding residue prediction with protein language models: does the structure matter?
abstract
MOTIVATION: Predicting protein-ligand binding sites is crucial in studying protein interactions with applications in biotechnology and drug discovery. Two distinct paradigms have emerged for this purpose: sequence-based methods, which leverage protein sequence information, and structure-based methods, which rely on the three-dimensional (3D) structure of the protein. Here, we analyze a hybrid approach that combines the strengths of both paradigms by integrating two recent deep learning architectures: protein language models (pLMs) from the sequence-based paradigm and Graph Neural Networks (GNNs) from the structure-based paradigm. Specifically, we construct a residue-level Graph Attention Network (GAT) model based on the protein's 3D structure that uses pre-trained pLM embeddings as node features. This integration enables us to study the interplay between the sequential information encoded in the protein sequence and the spatial relationships within the protein structure on the model performance. RESULTS: By exploiting a benchmark dataset over a range of ligands and ligand types, we have shown that using the structure information consistently enhances the predictive power of the baselines in absolute terms. Nevertheless, as more complex pLMs are used to represent node features, the relative impact of the structure information represented by the GNN architecture diminishes. The above observations suggest that although the use of the experimental protein structure almost always improves the accuracy of the prediction of the binding site, complex pLMs still contain structural information that leads to good predictive performance even without the use of 3D structure. AVAILABILITY AND IMPLEMENTATION: The datasets generated and/or analyzed during the current study, as well as pretrained models, are available in the following Zenodo link https://zenodo.org/records/15184302. The source code that was used to generate the results of the current study is available in the following GitHub repository https://github.com/hamzagamouh/pt-lm-gnn as well as in the following Zenodo link https://zenodo.org/records/15192327.
Hamza Gamouh, Marian Novotny, David Hoksza
Bioinform.3
2025 CryptoBench: cryptic protein-ligand binding sites dataset and benchmark
abstract
MOTIVATION: Structure-based methods for detecting protein-ligand binding sites play a crucial role in various domains, from fundamental research to biomedical applications. However, current prediction methodologies often rely on holo (ligand-bound) protein conformations for training and evaluation, overlooking the significance of the apo (ligand-free) states. This oversight is particularly problematic in the case of cryptic binding sites (CBSs) where holo-based assessment yields unrealistic performance expectations. RESULTS: To advance the development in this domain, we introduce CryptoBench, a benchmark dataset tailored for training and evaluating novel CBS prediction methodologies. CryptoBench is constructed upon a large collection of apo-holo protein pairs, grouped by UniProtID, clustered by sequence identity, and filtered to contain only structures with substantial structural change in the binding site. CryptoBench comprises 1107 structures with predefined cross-validation splits, making it the most extensive CBS dataset to date. To establish a performance baseline, we measured the predictive power of sequence- and structure-based CBS residue prediction methods using the benchmark. We selected PocketMiner as the state-of-the-art representative of the structure-based methods for CBS detection, and P2Rank, a widely-used structure-based method for general binding site prediction that is not specifically tailored for cryptic sites. For sequence-based approaches, we trained a neural network to classify binding residues using protein language model embeddings. Our sequence-based approach outperformed PocketMiner and P2Rank across key metrics, including area under the curve, area under the precision-recall curve, Matthew's correlation coefficient, and F1 scores. These results provide baseline benchmark results for future CBS and potentially also non-CBS prediction endeavors, leveraging CryptoBench as the foundational platform for further advancements in the field. AVAILABILITY AND IMPLEMENTATION: The CryptoBench dataset, including the benchmark model, is available on Open Science Framework-https://osf.io/pz4a9/. The code and tutorial are available at the GitHub repository-https://github.com/skrhakv/CryptoBench/.
Vít Skrhák, Marian Novotny, Christos P. Feidakis, Radoslav Krivák, David Hoksza
Bioinform.5
2024 Integrating Structural Features with Protein Language Models to Predict Protein-Ligand Binding Sites
abstract
Protein interactions are essential to biological function. Consequently, developing accurate models to predict ligand-binding sites is crucial not only for advancing our understanding of biological processes but also for applications like drug discovery. Advances in computational power have enabled machine learning techniques, including Protein Language Models (PLMs). In this study, we used the ESM-2 [1] PLM for binding site prediction and explored the integration of three-dimensional structural features to enhance its performance. We demonstrate that incorporating simple structural features can enhance the performance of sequence-based models. However, not all structural features are beneficial, indicating that some structural information is already embedded within the models themselves.
Matyás Brabec, David Hoksza
BIBM2
2024 Protein Family Sequence Generation through ProGen2 Fine-Tuning
abstract
Proteins are biomolecules involved in virtually all biological processes, making the design of novel proteins with specific functions crucial for advancing drug development and biological research. Large protein sequence databases allow for training language models adapted from natural language processing, treating amino acid sequences as a biological "language". However, these generative protein language models lack a straightforward, user-friendly method for prompting them to generate specific sequences with desired properties. In this work, we demonstrate how the pre-trained protein language model ProGen2 can be effectively fine-tuned for controllable generation of protein sequences from several distinct protein families. We validate the generated sequences using various in-silico metrics and show that the model is able to generate viable protein sequences that exhibit low similarity to existing proteins.
Hugo Hrbán, David Hoksza
BIBM2
2024 Protein Family Sequence Generation through ProGen2 Fine-Tuning
Hugo Hrbán, David Hoksza
BIBM2
2024 Protein Family Sequence Generation Through ProGen2 Fine-Tuning
abstract
Proteins are biomolecules involved in virtually all biological processes, making the design of novel proteins with specific functions crucial for advancing drug development and biological research. Large protein sequence databases allow for training language models adapted from natural language processing, treating amino acid sequences as a biological “language”. However, these generative protein language models lack a straightforward, user-friendly method for prompting them to generate specific sequences with desired properties. In this work, we demonstrate how the pre-trained protein language model ProGen2 can be effectively fine-tuned for controllable generation of protein sequences from several distinct protein families. We validate the generated sequences using various in-silico metrics and show that the model is able to generate viable protein sequences that exhibit low similarity to existing proteins.
Hugo Hrbán, David Hoksza
BIBM2
2024 Visualizations for universal deep-feature representations: survey and taxonomy
abstract
Abstract In data science and content-based retrieval, we find many domain-specific techniques that employ a data processing pipeline with two fundamental steps. First, data entities are represented by some visualizations, while in the second step, the visualizations are used with a machine learning model to extract deep features. Deep convolutional neural networks (DCNN) became the standard and reliable choice. The purpose of using DCNN is either a specific classification task or just a deep feature representation of visual data for additional processing (e.g., similarity search). Whereas the deep feature extraction is a domain-agnostic step in the pipeline (inference of an arbitrary visual input), the visualization design itself is domain-dependent and ad hoc for every use case. In this paper, we survey and analyze many instances of data visualizations used with deep learning models (mostly DCNN) for domain-specific tasks. Based on the analysis, we synthesize a taxonomy that provides a systematic overview of visualization techniques suitable for usage with the models. The aim of the taxonomy is to enable the future generalization of the visualization design process to become completely domain-agnostic, leading to the automation of the entire feature extraction pipeline. As the ultimate goal, such an automated pipeline could lead to universal deep feature data representations for content-based retrieval.
Tomás Skopal, Ladislav Peska, David Hoksza, Ivaná Sixtova, David Bernhauer
Knowl. Inf. Syst.3
2023 Framework for Protein Structures Conformation Analysis
abstract
Protein three-dimensional structure drives binding interactions, thus enabling proteins to perform their functions. However, a structure can undergo structural changes due to the changing environment. Analysis of datasets covering these structural states is challenging in terms of complexity and time consumption. Therefore, we propose a database framework designed to store, retrieve, and visualize information about changing structure conformations. The framework enables navigation between the series of transformations and investigation of each transformation individually via a web application. The framework sets the stage for a more accessible analysis of the dynamic molecular processes.
Vít Skrhák, David Hoksza
BIBM2
2023 Cryptic binding site prediction with protein language models
abstract
Structure-based identification of protein-ligand binding sites plays a crucial role in the initial stages of rational drug discovery pipelines. As machine learning methods are increasingly integrated into the process, a significant challenge arises while training these methods, as labeled data are typically derived from ligand-bound structures. Consequently, these methods struggle to detect binding sites within proteins where the binding site is concealed in the absence of a bound ligand. Here, we explore the possibility of harnessing protein language models to address this issue and compare their performance against state-of-the-art methods, both those specialized in the cryptic binding site (CBS) detection and those that are not. We show that applying pre-trained protein-language models in a relatively straightforward manner enables us to surpass the state-of-the-art of CBS prediction.
Vít Skrhák, Kamila Riedlova, Marian Novotny, David Hoksza
BIBM4
2023 Visual Representations for Data Analytics: User Study
Ladislav Peska, Ivaná Sixtova, David Hoksza, David Bernhauer, Tomás Skopal
CHIRA (2)3
2022 Exploration of protein sequence embeddings for protein-ligand binding site detection
abstract
Detection of protein-ligand binding sites is essential not only for protein function investigation but also in fields such as drug discovery or bioengineering. In this paper, we show that the recently-developed pre-trained language models can be used for protein-ligand binding site prediction. Specifically, we present a neural network architecture where inputs correspond to amino acids embeddings obtained from a protein language model. We show that increasing complexity of the language model improves the predictive performance of the method, eventually leading to results comparable to or surpassing state-of-the-art approaches. Unlike the existing methods, the presented approach does not require time-consuming computation of evolutionary information, resulting in faster running times.
David Hoksza, Hamza Gamouh
BIBM1
2022 Machine and human interpretable patient visualizations
abstract
A growing amount of data is stored in electronic health records, which are crucial for the clinical decision-making process. A large part of these data has a tabular form, consisting of numerical and categorical values originating from various laboratory examinations and sensors. Unlike medical images and clinical notes, tabular data lack higher semantics and, combined with the high dimensionality and heterogeneity, their interpretation by a human is challenging. On the other hand, we have witnessed superior performance of deep convolutional neural network (DCNN) models in the visual medical domain. In this paper, we propose visual representations of complex tabular medical data readable simultaneously by humans and machines. To show that these representations can encode the patient’s data semantics effectively, we use them to fine-tune a DCNN to predict the disability level of patients suffering from multiple sclerosis. Our experiments show that the visual models could match the performance of non-visual models. Moreover, the visual representations add the benefit o f s ummarizing complex information about the patient’s state to a human.
Ivaná Sixtova, Tomás Skopal, David Hoksza, Jakub Matejík, Tomás Uher
BIBM3
2022 AHoJ: rapid, tailored search and retrieval of apo and holo protein structures for user-defined ligands
abstract
SUMMARY: Understanding the mechanism of action of a protein or designing better ligands for it, often requires access to a bound (holo) and an unbound (apo) state of the protein. Resources for the quick and easy retrieval of such conformations are severely limited. Apo-Holo Juxtaposition (AHoJ), is a web application for retrieving apo-holo structure pairs for user-defined ligands. Given a query structure and one or more user-specified ligands, it retrieves all other structures of the same protein that feature the same binding site(s), aligns them, and examines the superimposed binding sites to determine whether each structure is apo or holo, in reference to the query. The resulting superimposed datasets of apo-holo pairs can be visualized and downloaded for further analysis. AHoJ accepts multiple input queries, allowing the creation of customized apo-holo datasets. AVAILABILITY AND IMPLEMENTATION: Freely available for non-commercial use at http://apoholo.cz. Source code available at https://github.com/cusbg/AHoJ-project. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Christos P. Feidakis, Radoslav Krivák, David Hoksza, Marian Novotny
Bioinform.3
2021 Closing the gap between formats for storing layout information in systems biology
abstract
The first version of this article listed one of its authors as Jan Hausenauer rather than Jan Hasenauer. This has now been corrected. The authors regret the error.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Jan Hasenauer, Reinhard Schneider 0002
Briefings Bioinform.1
2021 Machine Learning to Support the Presentation of Complex Pathway Graphs
abstract
Visualization of biological mechanisms by means of pathway graphs is necessary to better understand the often complex underlying system. Manual layout of such pathways or maps of knowledge is a difficult and time consuming process. Node duplication is a technique that makes layouts with improved readability possible by reducing edge crossings and shortening edge lengths in drawn diagrams. In this article, we propose an approach using Machine Learning (ML) to facilitate parts of this task by training a Support Vector Machine (SVM) with actions taken during manual biocuration. Our training input is a series of incremental snapshots of a diagram describing mechanisms of a disease, progressively curated by a human expert employing node duplication in the process. As a test of the trained SVM models, they are applied to a single large instance and 25 medium-sized instances of hand-curated biological pathways. Finally, in a user validation study, we compare the model predictions to the outcome of a node duplication questionnaire answered by users of biological pathways with varying experience. We successfully predicted nodes for duplication and emulated human choices, demonstrating that our approach can effectively learn human-like node duplication preferences to support curation of pathway diagrams in various contexts.
Sune S. Nielsen, Marek Ostaszewski, Fintan McGee, David Hoksza, Simone Zorzan
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 Closing the gap between formats for storing layout information in systems biology
abstract
The understanding of complex biological networks often relies on both a dedicated layout and a topology. Currently, there are three major competing layout-aware systems biology formats, but there are no software tools or software libraries supporting all of them. This complicates the management of molecular network layouts and hinders their reuse and extension. In this paper, we present a high-level overview of the layout formats in systems biology, focusing on their commonalities and differences, review their support in existing software tools, libraries and repositories and finally introduce a new conversion module within the MINERVA platform. The module is available via a REST API and offers, besides the ability to convert between layout-aware systems biology formats, the possibility to export layouts into several graphical formats. The module enables conversion of very large networks with thousands of elements, such as disease maps or metabolic reconstructions, rendering it widely applicable in systems biology.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Jan Hasenauer, Reinhard Schneider 0002
Briefings Bioinform.1
2019 MINERVA API and plugins: opening molecular network analysis and visualization to the community
abstract
SUMMARY: The complexity of molecular networks makes them difficult to navigate and interpret, creating a need for specialized software. MINERVA is a web platform for visualization, exploration and management of molecular networks. Here, we introduce an extension to MINERVA architecture that greatly facilitates the access and use of the stored molecular network data. It allows to incorporate such data in analytical pipelines via a programmatic access interface, and to extend the platform's visual exploration and analytics functionality via plugin architecture. This is possible for any molecular network hosted by the MINERVA platform encoded in well-recognized systems biology formats. To showcase the possibilities of the plugin architecture, we have developed several plugins extending the MINERVA core functionalities. In the article, we demonstrate the plugins for interactive tree traversal of molecular networks, for enrichment analysis and for mapping and visualization of known disease variants or known adverse drug reactions to molecules in the network. AVAILABILITY AND IMPLEMENTATION: Plugins developed and maintained by the MINERVA team are available under the AGPL v3 license at https://git-r3lab.uni.lu/minerva/plugins/. The MINERVA API and plugin documentation is available at https://minerva-web.lcsb.uni.lu.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Ewa Smula, Reinhard Schneider 0002
Bioinform.1
2018 Software framework for similarity-based prediction of protein interfaces
Jan Jelínek, Petr Skoda 0001, David Hoksza
BIBM3
2018 MolArt: a molecular structure annotation and visualization tool
abstract
Summary: MolArt fills the gap between sequence and structure visualization by providing a light-weight, interactive environment enabling exploration of sequence annotations in the context of available experimental or predicted protein structures. Provided a UniProt ID, MolArt downloads and displays sequence annotations, sequence-structure mapping and relevant structures. The sequence and structure views are interlinked, enabling sequence annotations being color overlaid over the mapped structures, thus providing an enhanced understanding and interpretation of the available molecular data. Availability and implementation: MolArt is released under the Apache 2 license and is available at https://github.com/davidhoksza/MolArt. The project web page https://davidhoksza.github.io/MolArt/ features examples and applications of the tool.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Reinhard Schneider 0002
Bioinform.1
2017 Template-based prediction of RNA tertiary structure using its predicted secondary structure
abstract
We present improved methodology for prediction of medium and large sized RNA three dimensional structures, based on comparative modeling approach. Method is enriched by the prediction of secondary structure, which is then used in tertiary structure prediction.
Rastislav Galvanek, David Hoksza
BIBM2
2017 Improving quality of ligand-binding site prediction with Bayesian optimization
abstract
Ligand binding site prediction from protein structure plays an important role in various complex rational drug design efforts. Its applications include drug side effects prediction, docking prioritization in inverse virtual screening and elucidation of protein function in genome wide structural studies. Currently available tools have limitations that disqualify them from many possible use cases. In general they are either fast and relatively inaccurate (e.g. purely geometric methods) or accurate but too slow for large scale applications (e.g. methods that rely on a large template libraries of known protein-ligand complexes). P2Rank is a recently introduced machine learning based method that have already exhibited speeds comparable to fastest geometric methods while providing much higher identification success rates. Here we present an improved version that brings speed-up as well as higher quality predictions. A leap in predictive performance was achieved thanks to the technique of Bayesian optimization, which allowed simultaneous optimization of numerous arbitrary parameters of the algorithm. We have evaluated our method with respect to various performance and prediction quality criteria and compared it to other state of the art methods, as well as to it's previous version, with encouraging results.
Radoslav Krivák, David Hoksza, Petr Skoda 0001
BIBM2
2017 Platform for ligand-based virtual screening integration
abstract
Ligand-based Virtual screening became a standard in-silico complement to the wet laboratory screening in small molecules discovery. It utilizes prior knowledge about molecules to rank a set of candidate molecules with yet unknown activity. Virtual screening software is often implemented as a specialized screening tool or as a part of a drug-discovery suite where ligand-based screening is one of the available tools. There exist many approaches to virtual screening, but although there exist several free-to-use screening tools, these tools usually provide only single approach to screening or do not provide user with aggregated analyses of multiple (preferably state-of-the-art) screening approaches. Such a functionality would be useful since different screening approaches yield different ranking of the candidate molecules. To the best of our knowledge there is no freely available, easy-to-use, tool that would allow to apply multiple screening approaches to a dataset and subsequently facilitate integration and interpretation of the results. Here, we present the first version of such a tool called ViSeT, available at https://github.com/skodapetr/viset under the MIT licence. ViSeT is a web-based, extensible, ligand-based virtual screening tool supporting multiple screening approaches and subsequent analysis of the results.
Petr Skoda 0001, David Hoksza, Jan Jelínek
BIBM2
2017 TRAVeLer: a tool for template-based RNA secondary structure visualization
abstract
BACKGROUND: Visualization of RNA secondary structures is a complex task, and, especially in the case of large RNA structures where the expected layout is largely habitual, the existing visualization tools often fail to produce suitable visualizations. This led us to the idea to use existing layouts as templates for the visualization of new RNAs similarly to how templates are used in homology-based structure prediction. RESULTS: This article introduces Traveler, a software tool enabling visualization of a target RNA secondary structure using an existing layout of a sufficiently similar RNA structure as a template. Traveler is based on an algorithm which converts the target and template structures into corresponding tree representations and utilizes tree edit distance coupled with layout modification operations to transform the template layout into the target one. Traveler thus accepts a pair of secondary structures and a template layout and outputs a layout for the target structure. CONCLUSIONS: Traveler is a command-line open source tool able to quickly generate layouts for even the largest RNA structures in the presence of a sufficiently similar layout. It is available at http://github.com/davidhoksza/traveler .
Richard Elias, David Hoksza
BMC Bioinform.2
2017 Utilizing knowledge base of amino acids structural neighborhoods to predict protein-protein interaction sites
abstract
BACKGROUND: Protein-protein interactions (PPI) play a key role in an investigation of various biochemical processes, and their identification is thus of great importance. Although computational prediction of which amino acids take part in a PPI has been an active field of research for some time, the quality of in-silico methods is still far from perfect. RESULTS: We have developed a novel prediction method called INSPiRE which benefits from a knowledge base built from data available in Protein Data Bank. All proteins involved in PPIs were converted into labeled graphs with nodes corresponding to amino acids and edges to pairs of neighboring amino acids. A structural neighborhood of each node was then encoded into a bit string and stored in the knowledge base. When predicting PPIs, INSPiRE labels amino acids of unknown proteins as interface or non-interface based on how often their structural neighborhood appears as interface or non-interface in the knowledge base. We evaluated INSPiRE's behavior with respect to different types and sizes of the structural neighborhood. Furthermore, we examined the suitability of several different features for labeling the nodes. Our evaluations showed that INSPiRE clearly outperforms existing methods with respect to Matthews correlation coefficient. CONCLUSION: In this paper we introduce a new knowledge-based method for identification of protein-protein interaction sites called INSPiRE. Its knowledge base utilizes structural patterns of known interaction sites in the Protein Data Bank which are then used for PPI prediction. Extensive experiments on several well-established datasets show that INSPiRE significantly surpasses existing PPI approaches.
Jan Jelínek, Petr Skoda 0001, David Hoksza
BMC Bioinform.3
2016 Template-based prediction of RNA tertiary structure
abstract
RNA tertiary structure prediction approaches can be divided into two groups: de novo methods and template-based modeling. De novo are applicable only for small molecules while in case of medium and large size RNA molecules, template-based modeling needs to be employed. While this type of modeling is quite common in protein structure prediction field, there exist only very few tools for template-based RNA structure prediction. Therefore, we present a methodology for prediction of RNA three dimensional structure (target) utilizing a known structure of a related RNA molecule (template). First, the target and template sequences are aligned. Next, sequentially similar regions in the alignment are identified and corresponding substructures are transferred from template to target. The remaining parts of the target structures are predicted using an external tool. This phase includes treatment of indels and valid linking of the transferred and predicted portions of the target structure. Our proposed method is able to predict even large ribosomal RNA structures when sufficiently similar template is available. The experiments have shown that the main impact on the quality of prediction has the sequence similarity of the template and target and number of indels. For structures with size of hundreds of nucleotides with sequence similarity with template over 50% and ratio of indels up to 50% the method is able to generate target structures up to ten RMSD with respect to the reference structure.
Rastislav Galvanek, David Hoksza, Josef Pánek
BIBM2
2016 Benchmarking platform for ligand-based virtual screening
abstract
Virtual screening (VS) of databases of chemical compounds has become a common step in the drug discovery process. Ligand-based virtual screening is a variant of VS where similarity to known active compounds is utilized in the discovery of new bioactive molecules. The cornerstone, which determines success of virtual screening, is the used molecular similarity measure. Currently, there is no superior approach to modeling molecular similarity and design of new similarity approaches is an active research field in cheminformatics. Therefore, proper benchmarking is of utter importance. In this paper, we describe common pitfalls of current approach to benchmarking of new methods. We focus on the importance of reproducibility and design of benchmarking datasets. Moreover, we identify the dataset difficulty as an important, yet not wildly utilized, property of the benchmarking data. To solve the identified issues we present a new benchmarking platform. The platform implements most commonly used molecular representations and includes datasets of varying difficulty levels as well as scripts which make the platform easy to use and extend. The existing representations are benchmarked using the proposed platform and results are presented. The benchmarking platform is available at https://github.com/skodapetr/lbvs-environment.
Petr Skoda 0001, David Hoksza
BIBM2
2016 Using Bayesian modeling on molecular fragments features for virtual screening
abstract
Virtual screening enables to search large small-molecule compound libraries for active molecules with respect to given macromolecular target. In ligand-based virtual screening, this goal is achieved by utilizing information about fragments or patterns present in existing known active compounds. Typically, the patterns are encoded as fingerprints which are used to screen a database of candidate compounds. In this work, we introduce an approach which uses Bayesian inference to encode activity-related information. Unlike previous approaches, our method does not utilize simple fragments, but rather uses features of these fragments. For each molecule, we generate a set of molecular fragments and extract molecular features for each of them. Next, we remove correlated features and use the remaining ones to build a Bayes model of activity. To score a previously unseen molecule, the molecule's fragment feature vectors are passed to the model and a score is obtained as the aggregation of their probability scores. When screening a database, this score is used to rank the compounds database. We show on datasets with various levels of difficulty that using fragments features rather then fragments themselves results in improvement of retrieval rates with respect to the best state-of-the art molecular fingerprints.
David Hoksza, Petr Skoda 0001
CIBCB1
2015 Exploration of topological torsion fingerprints
abstract
The screening of chemical libraries is an important step in identification of new leads in the drug discovery process. It is the size of the existing chemical libraries that renders laboratory screening expensive. A solution is to incorporate virtual screening into the process in order to reduce the number of molecules to be screened in the wet lab. In this paper, we explore several approaches to modification of one of the best performing methods for molecular representation in virtual screening campaigns, the topological torsion fingerprints. The modifications include the change of path length, altering atom descriptors and introduction of the so-called field version of the descriptors. With the field-based modification, our improved version of topological torsion fingerprints shows improvements by up to four percent in terms of area under the curve (AUC). The new topological torsion fingerprint thus represents one of the best performing molecular representation today.
Petr Skoda 0001, David Hoksza
BIBM2
2015 Activity-driven exploration of chemical space with morphing
abstract
Virtual screening (VS) methods, which became a common complement to the in vitro approaches in drug discovery projects, are naturally restricted by the compound libraries at hand while ignoring the wealth of compounds hidden in general chemical space. To close this gap, various methods for the exploration of chemical space have been proposed. One such approach is Molpher, a software framework that uses the technique of molecular morphing. Molecular morphing generates a series of compounds called morphs that represent a gradual structural transition between two given compounds. Because the exploration is driven solely by structural information, it disregards structurally diverse but possibly active compounds. Thus, we introduce the improvement of the algorithm where the exploration is driven by ligand biological activity rather than by its structure. On its input, the method takes a set of known active and inactive compounds. In the preparatory phase, feature selection is applied to choose descriptors that likely discriminate between active and inactive compounds. These features are then used to define a reference point towards which the exploration is directed. In the exploration phase, morphs are generated from all active compounds and Pareto-ranking scheme is applied to accept morphs for the next generation of molecular morphing. This iterative process results in structurally diverse molecules that share characteristic features of actives that separate them from inactives. The method was tested on four datasets from the PubChem BioAssay database. The results indicate that an activity-based exploration technique is able to generate structurally diverse compounds close to the selected point in the activity space. Thus, this technique is suitable for the generation of virtual libraries that can be further optimized and subsequently screened.
Martin Sícho, Daniel Svozil, David Hoksza
BIBM3
2015 MultiSETTER: web server for multiple RNA structure comparison
abstract
BACKGROUND: Understanding the architecture and function of RNA molecules requires methods for comparing and analyzing their tertiary and quaternary structures. While structural superposition of short RNAs is achievable in a reasonable time, large structures represent much bigger challenge. Therefore, we have developed a fast and accurate algorithm for RNA pairwise structure superposition called SETTER and implemented it in the SETTER web server. However, though biological relationships can be inferred by a pairwise structure alignment, key features preserved by evolution can be identified only from a multiple structure alignment. Thus, we extended the SETTER algorithm to the alignment of multiple RNA structures and developed the MultiSETTER algorithm. RESULTS: In this paper, we present the updated version of the SETTER web server that implements a user friendly interface to the MultiSETTER algorithm. The server accepts RNA structures either as the list of PDB IDs or as user-defined PDB files. After the superposition is computed, structures are visualized in 3D and several reports and statistics are generated. CONCLUSION: To the best of our knowledge, the MultiSETTER web server is the first publicly available tool for a multiple RNA structure alignment. The MultiSETTER server offers the visual inspection of an alignment in 3D space which may reveal structural and functional relationships not captured by other multiple alignment methods based either on a sequence or on secondary structure motifs.
Petr Cech, David Hoksza, Daniel Svozil
BMC Bioinform.2
2015 Multiple 3D RNA Structure Superposition Using Neighbor Joining
abstract
Recent advances in RNA research and the steady growth of available RNA structures call for bioinformatics methods for handling and analyzing RNA structural data. Recently, we introduced SETTER-a fast and accurate method for RNA pairwise structure alignment. In this paper, we describe MultiSETTER, SETTER extension for multiple RNA structure alignment. MultiSETTER combines SETTER's decomposition of RNA structures into non-overlapping structural subunits with the multiple sequence alignment algorithm ClustalW adapted for the structure alignment. The accuracy of MultiSETTER was assessed by the automatic classification of RNA structures and its comparison to SCOR annotations. In addition, MultiSETTER classification was also compared to multiple sequence alignment-based and secondary structure alignment-based classifications provided by LocARNA and RNADistance tools, respectively. MultiSETTER precompiled Windows libraries, as well as the C++ source code, are freely available from http://siret.cz/multisetter.
David Hoksza, Daniel Svozil
IEEE ACM Trans. Comput. Biol. Bioinform.1
2014 Scaffold-based chemical space exploration
abstract
The chemical space exploration is an in-silico lead discovery process which is not restricted by the existing compound libraries. On the other hand, the vastness of the chemical space can pose a limit on its application. That is also the case of a recently introduced molecular morphing-based method called Molpher which is focused on exploring the space between a pair of molecules by finding a connecting path between them. However, identification of this path is, in some cases, beyond the limits of the method due to the size of the space. Therefore, we are introducing a modified approach which utilizes chemical scaffolds in the exploration process. The new approach first simplifies the start and target compounds using their scaffolds, and subsequently finds a path within a much smaller space of scaffolds. This path forms a set of guides to be used for finding a path in the original chemical space. This way the originally complex problem is broken down into smaller problems which can be solved faster. Our method shows a significant speed-up over the existing approach (about 58%) and results in an increased number of cases in which the path is found for distant molecules.
David Hoksza, Petr Skoda 0001
BIBM1
2014 Template-based prediction of ribosomal RNA secondary structure
abstract
Determining the structure of ribosomal RNAs (rRNAs) is one of the crucial steps in understanding the process of protein synthesis, for which rRNAs are one of the basic components. Nevertheless, due to extreme technical difficulties, spatial (3D) structures have been resolved experimentally for only 14 organisms. Also, computational prediction of 3D rRNA structure is almost impossible, and prediction of secondary structure (the list of base pairs in the folded RNA), an important intermediate step between sequence and 3D structure that is used broadly in modeling of RNA structures, is in the case of rRNAs hindered by both extreme sequence length and high structure complexity. Here we present a proof-of-concept for an rRNA secondary structure prediction method that utilizes known structures as structural templates. Our template-based prediction algorithm determines those regions of the sequence for which structure is being predicted that are conserved well enough so that their secondary structure can be copied over from the template. The structure of the remaining, unconserved regions is predicted using a thermodynamic folding model. Applying a baseline implementation of our algorithm to the E. coli 16S rRNA, we have achieved state-of-the-art recall and precision using the structure of T. thermophilus 16S rRNA as a template.
Josef Pánek, Jan Hajic jr., David Hoksza
BIBM3
2014 2D Pharmacophore Query Generation
David Hoksza, Petr Skoda 0001
ISBRA1
2013 Chemical space visualization using ViFrame
abstract
Exploration of the chemical space is an important component of drug discovery process and its importance grows with the increase in the computation power which allows to explore larger areas of the chemical space. Recently, there emerged new algorithms proposed to automatically generate and search for compounds (objects in the chemical space) with desired properties. Although these approaches can be a big help, human interaction is usually still inevitable in the end. Visualization of the space can help make sense of the generated data and therefore visualization techniques are usually an integral part of any task related to chemical space exploration. Currently, there exist methods dealing with visualization of the chemical space but there is no framework supporting simple development of new methods. The purpose of this paper is to introduce such a modular framework called ViFrame. ViFrame offers the possibility to implement every single part of the visualization pipeline consisting of steps such as reading and merging molecules from multiple data sources, applying transformations and, of course, visualization of the data set in 2D space. The advantage of the framework consists in providing an environment where the user can focus on the development of the previously mentioned tasks while the framework supports seamless integration of the developed components. The framework also incorporates an application that provides the user with graphical interface for modules manipulation and presentation of the visualization results. For simple utilization of the application without the necessity of implementation of one's own module, several visualization methods have been implemented.
Petr Skoda 0001, David Hoksza
ICIS2
2012 On Optimizing the Non-metric Similarity Search in Tandem Mass Spectra by Clustering
David Hoksza, Jakub Lokoc, Tomás Skopal
ISBRA2
2012 SimTandem: Similarity Search in Tandem Mass Spectra
Jakub Galgonek, David Hoksza, Tomás Skopal
SISAP3
2012 Efficient RNA pairwise structure comparison by SETTER method
abstract
MOTIVATION: Understanding the architecture and function of RNA molecules requires methods for comparing and analyzing their 3D structures. Although a structural alignment of short RNAs is achievable in a reasonable amount of time, large structures represent much bigger challenge. However, the growth of the number of large RNAs deposited in the PDB database calls for the development of fast and accurate methods for analyzing their structures, as well as for rapid similarity searches in databases. RESULTS: In this article a novel algorithm for an RNA structural comparison SETTER (SEcondary sTructure-based TERtiary Structure Similarity Algorithm) is introduced. SETTER uses a pairwise comparison method based on 3D similarity of the so-called generalized secondary structure units. For each pair of structures, SETTER produces a distance score and an indication of its statistical significance. SETTER can be used both for the structural alignments of structures that are already known to be homologous, as well as for 3D structure similarity searches and functional annotation. The algorithm presented is both accurate and fast and does not impose limits on the size of aligned RNA structures. AVAILABILITY: The SETTER program, as well as all datasets, is freely available from http://siret.cz/hoksza/projects/setter/.
David Hoksza, Daniel Svozil
Bioinform.1
2011 Exploration of Chemical Space by Molecular Morphing
abstract
Many areas of chemical biology, such as drug discovery, rely heavily on chemical libraries offering compounds usable in the industrial processes. However, the "universe" containing all possible compounds, the so-called chemical space, is vast, and therefore, the libraries store only its representative parts. Thus, to explore the whole chemical space and to identify all its promising parts containing e.g., drug-like molecules computational methods have to be developed and employed. In this paper, we propose a method for traveling in the chemical space called Molpher. Given two molecules, Molpher is intended to find a sequence of related compounds, called path in the chemical space, leading from the starting molecule to the target one. The path is generated by iterative application of the so-called morphing operators corresponding to simple chemical operations such as adding or removing an atom or a bond. The molecules on the resulting path represent a focused library that can be used as a starting point for other experiments. We also propose a testbed for examining qualities of algorithms such as Molpher. The testbed is used to describe Molpher's qualities in terms of ability of finding a path in the space in the given time.
David Hoksza, Daniel Svozil
BIBE1
2011 SETTER - RNA SEcondary sTructure-based TERtiary Structure Similarity Algorithm
David Hoksza, Daniel Svozil
ISBRA1
2011 On metric approximations of the SProt measure
abstract
Protein structure similarity search is one of the most essential tasks in current proteomics. With the growing number of experimentally solved protein structures, the need for accurate and fast search methods is increasing. We have recently introduced similarity measure SProt showing good results in terms of effectiveness. However, this measure is computationally very intensive and, moreover, it is not suitable for metric indexing. In this paper, we introduce an efficient metric approximation of the SProt measure maintaining its former effectiveness.
Jakub Galgonek, David Hoksza
SISAP2
2011 Protein sequences identification using NM-tree
abstract
We have generalized a method for tandem mass spectra interpretation, based on the parameterized Hausdorff distance dHP. Instead of just peptides (short pieces of proteins), in this paper we describe the interpretation of whole protein sequences. For this purpose, we employ the recently introduced NM-tree to index the database of hypothetical mass spectra for exact or fast approximate search. The NM-tree combines the M-tree with the TriGen algorithm in a way that allows to dynamically control the retrieval precision at query time. A scheme for protein sequences identification using the NM-tree is proposed.
Tomás Skopal, David Hoksza, Jakub Lokoc, Jakub Galgonek
SISAP3
2009 DDPIn - Distance and density based protein indexing
abstract
Protein structure similarity and classification methods have many applications in protein function prediction and associated fields (e.g. drug discovery). In this paper, we propose a new protein structure representation method enabling fast and accurate classification. In our approach, each protein structure is represented by number of vectors (based on histogram of distances) equivalent to the number of its Calpharesidues. Each Calpharesidue represents a viewpoint from which the distances to each of the other residues are computed. Consequently, we use several methods to convert these distances into a n-dimensional feature vector which is indexed using a metric indexing structure (M-tree is the structure of our choice). While searching, we use single or multi-step approach which provides us with classification accuracy and speed comparable to the best contemporary classification methods.
David Hoksza
CIBCB1
2009 An application of the metric access methods to the mass spectrometry data
abstract
Mass spectrometry is a very popular method for protein and peptide identification nowadays. Abundance of data generated in this way grows exponentially every year. Although there exist algorithms for interpreting mass spectra, demand for faster and more accurate approaches remains. We propose an approach for preprocessing the protein sequence database based on metric access methods. This approach allows to select only a small set of suitable peptide sequence candidates, which can be then compared with experimental spectra using more sophisticated algorithms. We define logarithmic distance for selecting peptide sequence candidates and also outline possibilities of using the interval query for searching posttranslational modifications. The experimental results show that our approach is comparable in precision with nowadays most widely used public tools and outline possible directions for further research.
David Hoksza
CIBCB2
2008 Improved Alignment of Protein Sequences Based on Common Parts
David Hoksza
ISBRA1
2007 Improving the Performance of M-Tree Family by Nearest-Neighbor Graphs
Tomás Skopal, David Hoksza
ADBIS2
2007 Construction of Tree-Based Indexes for Level-Contiguous Buffering Support
Tomás Skopal, David Hoksza, Jaroslav Pokorný
DASFAA2