VLDB 2026 Research / reviewers in the wild / expert
Jirí Kléma
dblp:33/3665
· DBLP profile ↗
15ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-1753-9435ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-authorArtificial intelligence and machine learning · 4 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Utilizing RNA-seq data in monotone iterative generalized linear model to elevate prior knowledge quality of the circRNA-miRNA-mRNA regulatory axisabstractBACKGROUND: Current experimental data on RNA interactions remain limited, particularly for non-coding RNAs, many of which have only recently been discovered and operate within complex regulatory networks. Researchers often rely on in-silico interaction detection algorithms, such as TargetScan, which are based on biochemical sequence alignment. However, these algorithms have limited performance. RNA-seq expression data can provide valuable insights into regulatory networks, especially for understudied interactions such as circRNA-miRNA-mRNA. By integrating RNA-seq data with prior interaction networks obtained experimentally or through in-silico predictions, researchers can discover novel interactions, validate existing ones, and improve interaction prediction accuracy. RESULTS: This paper introduces Pi-GMIFS, an extension of the generalized monotone incremental forward stagewise (GMIFS) regression algorithm that incorporates prior knowledge. The algorithm first estimates prior response values through a prior-only regression, interpolates between these prior values and the original data, and then applies the GMIFS method. Our experimental results on circRNA-miRNA-mRNA regulatory interaction networks demonstrate that Pi-GMIFS consistently enhances precision and recall in RNA interaction prediction by leveraging implicit information from bulk RNA-seq expression data, outperforming the initial prior knowledge. CONCLUSION: Pi-GMIFS is a robust algorithm for inferring acyclic interaction networks when the variable ordering is known. Its effectiveness was confirmed through extensive experimental validation. We proved that RNA-seq data of a representative size help infer previously unknown interactions available in TarBase v9 and improve the quality of circRNA disease annotation. Alikhan Anuarbekov, Jirí Kléma |
BMC Bioinform. | 2 |
| 2024 | Directed Evolution of Proteins via Bayesian Optimization in Embedding SpaceabstractDirected evolution is an iterative laboratory process of designing proteins with improved function by iteratively synthesizing new protein variants and evaluating their desired property with expensive and time-consuming biochemical screening. Machine learning methods can help select informative or promising variants for screening to increase their quality and reduce the amount of necessary screening. In this paper, we present a novel method for machine-learning-assisted directed evolution of proteins which combines Bayesian optimization with informative representation of protein variants extracted from a pre-trained protein language model. We demonstrate that the new representation based on the sequence embeddings significantly improves the performance of Bayesian optimization yielding better results with the same number of conducted screening in total. At the same time, our method outperforms the state-of-the-art machine-learning-assisted directed evolution methods with regression objective. Matous Soldát, Jirí Kléma |
BIBM | 2 |
| 2022 | On functional annotation with gene co-expression networksabstractGene co-expression networks have frequently been used for functional annotation. In these networks, an unknown gene is annotated with terms that have already been associated with genes whose expression profiles t end to correlate with the expression profile of the unknown gene. Despite the biological plausibility of this principle referred to as guilt-by-association, its applicability has not been thoroughly experimentally verified yet. In our paper, we formulate several statistical hypotheses concerning the principle and test them on a representative expression dataset. We demonstrate that gene annotation carried out with co-expression networks clearly outperforms random annotation and improves with increasing sample size and the knowledge of gene co-location. Eventually, we discuss the practical significance of this way of functional annotation. Vladimír Kunc, Jirí Kléma |
BIBM | 2 |
| 2022 | circGPA: circRNA functional annotation based on probability-generating functionsabstractRecent research has already shown that circular RNAs (circRNAs) are functional in gene expression regulation and potentially related to diseases. Due to their stability, circRNAs can also be used as biomarkers for diagnosis. However, the function of most circRNAs remains unknown, and it is expensive and time-consuming to discover it through biological experiments. In this paper, we predict circRNA annotations from the knowledge of their interaction with miRNAs and subsequent miRNA-mRNA interactions. First, we construct an interaction network for a target circRNA and secondly spread the information from the network nodes with the known function to the root circRNA node. This idea itself is not new; our main contribution lies in proposing an efficient and exact deterministic procedure based on the principle of probability-generating functions to calculate the p-value of association test between a circRNA and an annotation term. We show that our publicly available algorithm is both more effective and efficient than the commonly used Monte-Carlo sampling approach that may suffer from difficult quantification of sampling convergence and subsequent sampling inefficiency. We experimentally demonstrate that the new approach is two orders of magnitude faster than the Monte-Carlo sampling, which makes summary annotation of large circRNA files feasible; this includes their reannotation after periodical interaction network updates, for example. We provide a summary annotation of a current circRNA database as one of our outputs. The proposed algorithm could be generalized towards other types of RNA in way that is straightforward. Petr Rysavý, Jirí Kléma, Michaela Dostálová Merkerová |
BMC Bioinform. | 2 |
| 2015 | Sparse omics-network regularization to increase interpretability and performance of linear classification modelsabstractCurrent high-throughput technologies lead to the boost of omics data with thousands of features measured in parallel. The phenotype specific markers are learned from the data to better understand the disease mechanism and to build predictive models. However, the learning is prone to overfitting, caused by a small sample size and large feature space dimension. Consequently, resulting models are inaccurate and difficult to interpret due to the complex nature of omics processes. In this paper, we propose a methodology for learning simple yet biologically meaningful linear classification models. A linear support vector machine is trained; the learning is regularized by prior knowledge. Regularization parameters enable the expert to operatively adjust the interpretation of the models and their conformity with recent domain research while maintaining their accuracy. We performed robust experiments showing empirical validity of our methodology. In the study related to myelodysplastic syndrome we demonstrate the performance and interpretation of disease classification models. These models are consistent with recent progress in myelodysplastic syndrome research. Michael Andel, Filip Masri, Jirí Kléma, Zdenek Krejcík, Monika Belickova |
BIBM | 3 |
| 2014 | Network-constrained forest for regularized omics data classificationabstractContemporary molecular biology deals with a wide and heterogeneous set of measurements to model and understand underlying biological processes including complex diseases. Machine learning provides a frequent approach to build such models. However, the models built solely from measured data often suffer from overfitting, as the sample size is typically much smaller than the number of measured features. In this paper, we propose a random forest-based classifier that minimizes this overfitting with the aid of prior knowledge in the form of a feature interaction network. We illustrate the proposed method in the task of disease classification based on measured mRNA and miRNA profiles complemented by the interaction network composed of the miRNA-mRNA target relations and mRNA-mRNA interactions corresponding to the interactions between their encoded proteins. We demonstrate that the proposed network-constrained forest employs prior knowledge to increase learning bias and consequently to improve classification accuracy, stability and comprehensibility of the resulting model. The experiments are carried out in the domain of myelodysplastic syndrome that we are concerned about in the long term. We validate our approach in the public domain of ovarian carcinoma, with the same data form. We believe that the idea of a network-constrained forest can straightforwardly be generalized towards arbitrary omics data with an available and non-trivial feature interaction network. Michael Andel, Jirí Kléma, Zdenek Krejcík |
BIBM | 2 |
| 2014 | miXGENE Tool for Learning from Heterogeneous Gene Expression Data Using Prior KnowledgeabstractHigh-throughput genomic technologies have proved to be useful in the search for both genetic disease markers and more complex predictive and descriptive models. By the same token, it became obvious that accurate and interpretable models need to concern more than raw measurements taken at a single phase of gene expression. In order to reach a deeper understanding of the molecular nature of complexly orchestrated biological processes, all the available measurements and existing genomic knowledge need to be fused. In this paper, we introduce a tool for machine learning from heterogeneous gene expression data using prior knowledge. The tool is called miXGENE, it is elaborated upon in close connection with the biological departments that dispose of the above-mentioned data and have a strong interest in their integration within particular problem-oriented projects. The main idea is not merely to capture the transcriptional phase of gene expression quantified by the amount of messenger RNA (mRNA). The increasing availability of microRNA (miRNA) data asks for its concurrent analysis with the transcriptional data. Moreover, epigenetic data such as methylation measurements can help to explain unexpected transcriptional irregularities. miXGENE is an environment for building workflows that enable rapid prototyping of integrative molecular models. Matej Holec, Valentin Gologuzov, Jirí Kléma |
CBMS | 3 |
| 2012 | Comparative evaluation of set-level techniques in predictive classification of gene expression samplesabstractBACKGROUND: Analysis of gene expression data in terms of a priori-defined gene sets has recently received significant attention as this approach typically yields more compact and interpretable results than those produced by traditional methods that rely on individual genes. The set-level strategy can also be adopted with similar benefits in predictive classification tasks accomplished with machine learning algorithms. Initial studies into the predictive performance of set-level classifiers have yielded rather controversial results. The goal of this study is to provide a more conclusive evaluation by testing various components of the set-level framework within a large collection of machine learning experiments. RESULTS: Genuine curated gene sets constitute better features for classification than sets assembled without biological relevance. For identifying the best gene sets for classification, the Global test outperforms the gene-set methods GSEA and SAM-GS as well as two generic feature selection methods. To aggregate expressions of genes into a feature value, the singular value decomposition (SVD) method as well as the SetSig technique improve on simple arithmetic averaging. Set-level classifiers learned with 10 features constituted by the Global test slightly outperform baseline gene-level classifiers learned with all original data features although they are slightly less accurate than gene-level classifiers learned with a prior feature-selection step. CONCLUSION: Set-level classifiers do not boost predictive accuracy, however, they do achieve competitive accuracy if learned with the right combination of ingredients. AVAILABILITY: Open-source, publicly available software was used for classifier learning and testing. The gene expression datasets and the gene set database used are also publicly available. The full tabulation of experimental results is available at http://ida.felk.cvut.cz/CESLT. Matej Holec, Jirí Kléma, Filip Zelezný, Jakub Tolar |
BMC Bioinform. | 2 |
| 2012 | Empirical Evidence of the Applicability of Functional Clustering through Gene Expression ClassificationabstractThe availability of a great range of prior biological knowledge about the roles and functions of genes and gene-gene interactions allows us to simplify the analysis of gene expression data to make it more robust, compact, and interpretable. Here, we objectively analyze the applicability of functional clustering for the identification of groups of functionally related genes. The analysis is performed in terms of gene expression classification and uses predictive accuracy as an unbiased performance measure. Features of biological samples that originally corresponded to genes are replaced by features that correspond to the centroids of the gene clusters and are then used for classifier learning. Using 10 benchmark data sets, we demonstrate that functional clustering significantly outperforms random clustering without biological relevance. We also show that functional clustering performs comparably to gene expression clustering, which groups genes according to the similarity of their expression profiles. Finally, the suitability of functional clustering as a feature extraction technique is evaluated and discussed. Milos Krejnik, Jirí Kléma |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | Comparative Evaluation of Set-Level Techniques in Microarray Classification
Jirí Kléma, Matej Holec, Filip Zelezný, Jakub Tolar |
ISBRA | 1 |
| 2009 | Integrating Multiple-Platform Expression Data through Gene Set Features
Matej Holec, Filip Zelezný, Jirí Kléma, Jakub Tolar |
ISBRA | 3 |
| 2008 | Sequential Data Mining: A Comparative Case Study in Development of Atherosclerosis Risk FactorsabstractSequential data represent an important source of potentially new medical knowledge. However, this type of data is rarely provided in a format suitable for immediate application of conventional mining algorithms. This paper summarizes and compares three different sequential mining approaches based, respectively, on windowing, episode rules, and inductive logic programming. Windowing is one of the essential methods of data preprocessing. Episode rules represent general sequential mining, while inductive logic programming extracts first-order features whose structure is determined by background knowledge. The three approaches are demonstrated and evaluated in terms of a case study STULONG. It is a longitudinal preventive study of atherosclerosis where the data consist of a series of long-term observations recording the development of risk factors and associated conditions. The intention is to identify frequent sequential/temporal patterns. Possible relations between the patterns and an onset of any of the observed cardiovascular diseases are also studied. Jirí Kléma, Lenka Nováková, Filip Karel, Olga Stepánková, Filip Zelezný |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2006 | Mining Plausible Patterns from Genomic DataabstractThe discovery of biologically interpretable knowledge from gene expression data is one of the largest contemporary genomic challenges. As large volumes of expression data are being generated, there is a great need for automated tools that provide the means to analyze them. However, the same tools can provide an overwhelming number of candidate hypotheses which can hardly be manually exploited by an expert. An additional knowledge helping to focus automatically on the most plausible candidates only can up-value the experiment significantly. Background knowledge available in literature databases, biological ontologies and other sources can be used for this purpose. In this paper we propose and verify a methodology that enables to effectively mine and represent meaningful over-expression patterns. Each pattern represents a bi-set of a gene group over-expressed in a set of biological situations. The originality of the framework consists in its constraint-based nature and an effective cross-fertilization of constraints based on expression data and background knowledge. The result is a limited set of candidate patterns that are most likely interpretable by biologists. Supplemental automatic interpretations serve to ease this process. Various constraints can generate plausible pattern sets of different characteristics. Jirí Kléma, Arnaud Soulet, Bruno Crémilleux, Sylvain Blachon, Olivier Gandrillon |
CBMS | 1 |
| 2002 | Optimized Model Tuning in Medical SystemsabstractFor patients considering elective major surgery, information about operative mortality risks is essential for careful decision making. To help patients and surgeons make informed decisions about whether to undergo elective high-risk surgery, a reliable predictive model would be beneficial. This paper focuses on development and optimized tuning of a model predicting risks related to heart interventions of several types. The model is based oil representative data sets collected in the Merged National Registry (MNR) on Cardiovascular Interventions. The registry is operated and governed by the MEDICON Center. The central attention is paid to an instance-based reasoning model and its tuning. In particular, the paper presents and discusses benefits of utilizing a genetic algorithm with limited convergence for this purpose. Jirí Kléma, Jirí Kubalík, Jiri Palous |
CBMS | 1 |
| 2001 | Evolving Groups of Basic Decision TreesabstractA decision tree is a good classifier with a transparent decision mechanism. Decision-tree building methods usually have problems in splitting the learning samples into more subsets, because of the nature of the tree. If the classification into such subsets is not possible, it is better to put the classification decision on to some other classifier. This leads to the introduction of a null classification, which simply means that no classification is possible in this step. This approach is sensible with evolutionary methods, as they can handle a number of trees simultaneously. In the process of construction, we have to address the problem of whether a classification is sensible. The performance of the proposed model has been tested on several data sets and the results presented on one such data set show its potential. Matej Sprogar, Peter Kokol, Milan Zorman, Vili Podgorelec, Lenka Lhotská, Jirí Kléma |
CBMS | 6 |